{"id":124893,"date":"2026-07-30T21:46:08","date_gmt":"2026-07-30T21:46:08","guid":{"rendered":"https:\/\/www.europesays.com\/ai\/124893\/"},"modified":"2026-07-30T21:46:08","modified_gmt":"2026-07-30T21:46:08","slug":"ai-probability-bias-traced-to-training-claude-pessimistic-gpt-optimistic-across-all-tests","status":"publish","type":"post","link":"https:\/\/www.europesays.com\/ai\/124893\/","title":{"rendered":"AI Probability Bias Traced to Training: Claude Pessimistic, GPT Optimistic Across All Tests"},"content":{"rendered":"<p>Ask an AI to estimate a startup&#8217;s odds of success. Now ask the same AI to estimate that startup&#8217;s odds of failure. If the two numbers add up to anything other than 100%, you&#8217;ve found a distortion that standard AI benchmarks were never built to catch \u2014 and according to research posted yesterday to arXiv, that distortion is systematic, directional, and traceable to one specific stage of how frontier AI models are built.<\/p>\n<p>The paper, which introduced a new measurement framework called OptimismBench, <a href=\"https:\/\/arxiv.org\/abs\/2607.26981\" rel=\"nofollow noopener\" target=\"_blank\">tested 16 models from eight major AI providers<\/a> and found that 14 of them consistently overestimate the probability of positive outcomes. The two exceptions are both from Anthropic: Claude Sonnet 4.6 and Claude Opus 4.6 lean in the opposite direction, systematically underestimating success probabilities relative to what <a href=\"https:\/\/arxiv.org\/abs\/2607.26981\" rel=\"nofollow noopener\" target=\"_blank\">the math should allow<\/a>. The authors \u2014 Seonglae Cho and Adriano Koshiyama of Holistic AI and University College London \u2014 traced the divergence to a single cause: the alignment choices made after a model&#8217;s initial pretraining. Which direction a model errs in its probability estimates is, in their finding, largely a fingerprint of the lab that built it.<\/p>\n<p>The researchers illustrate this gap with a concrete example: an AI that rates a startup&#8217;s chance of success at 70% but its chance of failure at only 15% is <a href=\"https:\/\/arxiv.org\/abs\/2607.26981\" rel=\"nofollow noopener\" target=\"_blank\">not merely imprecise<\/a>. The two estimates should sum to approximately 100%. The missing 15 percentage points are a directional distortion \u2014 and no existing aggregate calibration benchmark would flag it, because aggregate metrics do not track direction. For any domain in which AI is used as a decision aid \u2014 risk assessment, investment screening, project timeline forecasting, medical outcome estimation \u2014 this is not a minor statistical quirk. It is a systematic thumb on the scale.<\/p>\n<p>Why Standard Calibration Benchmarks Miss the Problem<\/p>\n<p>The AI field has well-established tools for measuring whether a model&#8217;s probability estimates match reality \u2014 a property researchers call calibration. Standard calibration metrics, such as Expected Calibration Error, measure how far a model&#8217;s predictions are from the true frequency of outcomes over large samples. The problem OptimismBench identifies is not that these metrics are wrong. It is that they aggregate unsigned errors: overestimates and underestimates cancel each other out in the aggregate score, so a model that consistently overestimates success by 15 percentage points and consistently underestimates failure by 15 percentage points could record a perfect calibration score \u2014 while delivering systematically misleading answers in every individual use.<\/p>\n<p>A second structural limitation makes the problem harder to detect: many real-world probability questions have no ground-truth answer to measure against. You cannot compute calibration error on a startup outcome before it happens. A benchmark that requires ground truth cannot test models on the kinds of probabilistic judgment people actually use AI for most. OptimismBench&#8217;s detection method is designed around this constraint.<\/p>\n<p>How OptimismBench Detects Directional Bias Without Ground Truth<\/p>\n<p>The benchmark&#8217;s core mechanism is the inverted pair. For each of <a href=\"https:\/\/arxiv.org\/abs\/2607.26981\" rel=\"nofollow noopener\" target=\"_blank\">60 scenarios<\/a>, the model is asked to estimate both the probability of success and the probability of failure for the same event. Logically, those two estimates should sum to 100%. A model that consistently produces estimates where P(success) &gt; 100% \u2212 P(failure) is systematically optimistic; one where P(success) &lt; 100% \u2212 P(failure) is systematically pessimistic. The signed difference \u2014 the researchers call it the &#8220;Skew&#8221; score \u2014 quantifies the directional distortion without needing any external outcome data.<\/p>\n<p>The full benchmark spans <a href=\"https:\/\/arxiv.org\/abs\/2607.26981\" rel=\"nofollow noopener\" target=\"_blank\">3,870 items across 10 languages<\/a> and covers four axiom batteries: conjunction fallacies, conditional probability, dose-response monotonicity, and explicit calibration items. The multilingual scope is not incidental. A secondary finding is that the language in which you pose a probability question to an AI makes almost no difference to the direction of its error: <a href=\"https:\/\/arxiv.org\/abs\/2607.26981\" rel=\"nofollow noopener\" target=\"_blank\">inter-model variance in Skew scores was 4.7 times larger than inter-language variance<\/a>. Whether you ask GPT for a probability estimate in French, Japanese, or English, the directional tilt does not meaningfully change. Which lab built the model is the variable that matters.<\/p>\n<p>Sign by Lab, Magnitude by Scale: The Full Results<\/p>\n<p>Testing 16 models from eight providers \u2014 among them OpenAI&#8217;s GPT family, Google&#8217;s Gemini, Meta&#8217;s Llama variants, Mistral, DeepSeek, Zhipu&#8217;s GLM series, and Anthropic&#8217;s Claude \u2014 the researchers found that <a href=\"https:\/\/arxiv.org\/abs\/2607.26981\" rel=\"nofollow noopener\" target=\"_blank\">all 16 show statistically significant directional bias<\/a>. None are neutral.<\/p>\n<p>The pattern the researchers describe as &#8220;sign by lab, magnitude by scale&#8221; has two components. First: which direction a model errs is largely determined by the lab that built it. Among the 16 models tested, 14 skew optimistic. The two outliers are exclusively from Anthropic, whose Claude Sonnet 4.6 and Claude Opus 4.6 produce pessimistic bias \u2014 consistently underestimating success probabilities relative to what their failure-probability estimates imply. Second: how large the error is scales with model size within a given lab&#8217;s family.<\/p>\n<p>An internal complication within Anthropic&#8217;s results is worth noting. Claude Haiku 4.5, the smallest model in the family, shows a <a href=\"https:\/\/arxiv.org\/abs\/2607.26981\" rel=\"nofollow noopener\" target=\"_blank\">positive (optimistic) Skew score of +6.1<\/a>, placing it in the same direction as the majority of other models tested. This means the pessimistic pattern is not a blanket property of &#8220;Anthropic models&#8221; but appears tied to scale within the family: the larger frontier models \u2014 Sonnet 4.6 and Opus 4.6 \u2014 are the outliers, while the smaller Haiku follows the industry-wide pattern.<\/p>\n<p>Post-Training Is What Installs the Bias<\/p>\n<p>The most consequential finding in the paper \u2014 and the one that distinguishes it from earlier work on AI overconfidence \u2014 is the mechanism. By comparing <a href=\"https:\/\/arxiv.org\/abs\/2607.26981\" rel=\"nofollow noopener\" target=\"_blank\">11 matched pairs of base models against their instruction-tuned or chat-aligned counterparts<\/a> across four model families, the team was able to observe directly how directional bias changes during the post-training phase \u2014 the stage of AI development where models are fine-tuned using alignment techniques.<\/p>\n<p>Post-training does not merely amplify a bias that was already there. It sets the sign of the directional distortion. Even more striking: the direction of the shift differs between model families \u2014 in some families, post-training pushes the model toward optimism; in Anthropic&#8217;s frontier tier, it pushes in the opposite direction.<\/p>\n<p>The training approach is the most plausible explanation for that divergence. Reinforcement learning from human feedback \u2014 the dominant post-training technique across most major AI labs \u2014 is <a href=\"https:\/\/arxiv.org\/abs\/2602.01002\" rel=\"nofollow noopener\" target=\"_blank\">now well-documented to amplify sycophantic behavior<\/a> because human annotators systematically prefer responses that are encouraging and positively framed. A model trained to produce human-preferred outputs learns to frame things optimistically because optimism is preferred. That is the structural mechanism through which RLHF installs an optimistic directional tilt. Anthropic&#8217;s <a href=\"https:\/\/www.anthropic.com\/research\/constitutional-ai-harmlessness-from-ai-feedback\" rel=\"nofollow noopener\" target=\"_blank\">Constitutional AI approach<\/a> uses principle-based feedback rather than raw human approval ratings \u2014 which does not have this structural bias toward positive framing. The pessimistic tilt in Anthropic&#8217;s frontier models is, on this account, what you would expect from a training methodology that replaced &#8220;does the annotator like this?&#8221; with &#8220;does this response adhere to stated principles?&#8221;<\/p>\n<p>This connection between the post-training mechanism and the optimism bias direction is the largest implication the research points toward. The sycophancy research and the probability-bias research are likely measuring the same underlying phenomenon from two different angles: RLHF rewards agreement and encouragement; agreement and encouragement are directionally optimistic; therefore RLHF-aligned models skew their probability estimates in the optimistic direction. The alignment choice is the bias-installation mechanism.<\/p>\n<p>Independent corroboration of this pattern comes from separate research on agentic AI performance. A study on budget estimation across agentic tasks found that <a href=\"https:\/\/arxiv.org\/abs\/2606.00198\" rel=\"nofollow noopener\" target=\"_blank\">Claude Sonnet and Opus were &#8220;closest to calibrated, but still skew low&#8221;<\/a> on their predictions about remaining task budget, while Gemini and Qwen were the most optimistic \u2014 a directional alignment consistent with OptimismBench&#8217;s findings from a completely different measurement approach.<\/p>\n<p>What &#8220;Pessimistic&#8221; Actually Means for Frontier Claude Users<\/p>\n<p>The finding that Anthropic&#8217;s frontier models skew pessimistic in probability estimates does not straightforwardly mean Claude is better or worse for decision-making \u2014 it means it is differently biased. In contexts where the risk of overconfidence is the primary concern (investment screening, medical risk assessment, project timeline planning where optimistic estimates lead to cascading failures), a model that underestimates success probabilities may be the more conservative and appropriate tool. In contexts where motivation or goal-setting is the concern, a pessimistic probability estimate may be counterproductive.<\/p>\n<p>The more direct implication for any user who employs Claude for probability-sensitive decisions is: the estimates you receive from Claude Sonnet 4.6 or Claude Opus 4.6 will tend to understate success probabilities relative to the math implied by the complementary failure probability. The estimates you receive from GPT-family models will tend to overstate them. Neither is neutral. The direction of the error is a product of how the model was trained, not of the evidence it was given.<\/p>\n<p>The paper tested four specific interventions to see whether the bias could be easily corrected: varied prompt wording, different temperature settings, narrative perspective shifts (posing the question from a worried investor&#8217;s perspective versus an enthusiastic founder&#8217;s), and explicit self-debiasing instructions that asked the model to check its own estimates for systematic bias. None of these interventions <a href=\"https:\/\/arxiv.org\/abs\/2607.26981\" rel=\"nofollow noopener\" target=\"_blank\">changed the direction of the Skew score<\/a>. They affected only the magnitude. The bias is not a surface artifact of prompting style. It is embedded in the model&#8217;s probability structure.<\/p>\n<p>Does AI Language Still Matter? The Multilingual Finding<\/p>\n<p>One of the paper&#8217;s secondary findings has direct practical relevance for anyone who uses AI in multiple languages. A 17-model, six-language comparison found that a model&#8217;s identity \u2014 the lab that built it \u2014 is a substantially stronger predictor of directional probability bias than the language in which the query is posed. <a href=\"https:\/\/arxiv.org\/abs\/2607.26981\" rel=\"nofollow noopener\" target=\"_blank\">Inter-model variance in Skew scores was 4.7 times the inter-language variance<\/a>.<\/p>\n<p>This is a meaningful qualification to the recently documented finding that Claude&#8217;s expressed behavioral values shift by language \u2014 with Hindi producing the most validating responses and English\/Russian the most rigorous. That finding concerned the affective and value-expressive dimensions of Claude&#8217;s behavior (how warm or rigorous it sounds). The OptimismBench finding concerns the probability-calibration dimension of Claude&#8217;s behavior (how it estimates numerical likelihoods). The two dimensions are related \u2014 sycophancy and optimism bias likely share a training-mechanism root \u2014 but they are not identical. Language strongly affects the affective dimension of AI behavior; it barely affects the directional probability dimension. The probability bias is a more stable, lab-determined property.<\/p>\n<p>What This Benchmark Makes Possible<\/p>\n<p>For AI developers and evaluators, OptimismBench provides a concrete diagnostic that requires no ground-truth outcome data and no lab-specific benchmarking infrastructure. Running the inverted-pair test on any model \u2014 asking P(success) and P(failure) for the same scenarios and computing the signed average deviation \u2014 immediately surfaces whether alignment choices have installed a directional thumb on the probability scale, and in which direction.<\/p>\n<p>For organizations deploying AI in any context that involves probability assessment, the research is a reminder that the reported confidence level is a function of training as much as of evidence. A model that says &#8220;70% chance of success&#8221; has produced an output shaped partly by which AI company trained it and what that company&#8217;s alignment approach systematically rewards.<\/p>\n<p>The full 3,870-item OptimismBench dataset has been publicly released for per-model directional-bias auditing. The paper is currently under review at EMNLP 2026.<\/p>\n<p>Frequently Asked QuestionsWhy do most AI models overestimate the probability of success?<\/p>\n<p>The most likely mechanism is reinforcement learning from human feedback (RLHF), the dominant post-training technique used by most major AI labs. Human annotators consistently prefer responses that are encouraging, affirming, and positively framed. A model trained to maximize human approval ratings learns to frame outcomes optimistically \u2014 and this manifests as a systematic tilt toward overestimating positive probabilities. The OptimismBench finding that post-training sets the sign of the directional bias, and that different alignment approaches produce opposite directions, is consistent with this mechanism: labs that use conventional RLHF produce optimistic models; Anthropic&#8217;s Constitutional AI approach, which uses principle-based rather than human-approval-based feedback, produces pessimistic frontier models instead.<\/p>\n<p>Is Claude&#8217;s pessimism better or worse than other AI models&#8217; optimism for real decisions?<\/p>\n<p>It depends on the decision. For risk assessment and downside planning \u2014 scenarios where overconfidence is the greater danger \u2014 a model that underestimates success probabilities may produce more conservative and appropriate guidance. For motivation, goal-setting, or contexts where accurate upside probability matters, pessimistic estimates could lead to underinvestment or avoidance of good opportunities. Neither direction is categorically better; both are systematic errors. The practical implication is that Claude&#8217;s frontier-tier probability estimates for success scenarios are likely understated relative to the math implied by its complementary failure estimates \u2014 and that prompt-rewording, perspective-shifting, or asking the model to check its own work will not correct this. The bias is embedded in the model&#8217;s probability structure, not in its prompting surface.<\/p>\n<p>Can I use OptimismBench to test any AI model I use at work?<\/p>\n<p>The 3,870-item dataset has been publicly released for per-model auditing. Running the core test does not require statistical expertise: for any scenario, ask the model to estimate both the probability of success and the probability of failure, and check whether the two numbers sum to approximately 100%. If they systematically diverge \u2014 success estimates consistently higher than 100% minus the failure estimate \u2014 the model is optimistically biased in that domain. If they diverge in the other direction, it is pessimistically biased. This inverted-pair check works without needing any outcome data and can be applied to any probability-sensitive domain where you use AI as a decision aid.<\/p>\n<p>What does the direction of an AI&#8217;s probability bias tell me about how it was trained?<\/p>\n<p>According to this research, quite a lot. The finding that post-training sets the sign of the directional bias, with different model families shifting in opposite directions, means the bias direction is a detectable fingerprint of a lab&#8217;s alignment methodology. Models trained primarily through human preference optimization (conventional RLHF) appear to acquire optimistic bias because human annotators favor positive-framed, encouraging responses. Anthropic&#8217;s frontier models \u2014 which use Constitutional AI, a principle-based rather than human-approval-based alignment approach \u2014 skew pessimistic instead. This does not mean pessimism is the goal; it is the directional artifact of not reinforcing human approval of positive framings. The research suggests that any organization evaluating AI models for high-stakes probability estimation should treat bias direction alongside accuracy as a required disclosure \u2014 because the direction tells you something specific about the training process that produced the model.<\/p>\n<p>&#8220;OptimismBench: Forecasting Bias and the Alignment Effect in Language Model Judgment&#8221; was posted to arXiv on July 29, 2026 (arXiv:2607.26981). Authors: Seonglae Cho and Adriano Koshiyama, Holistic AI and University College London. The 3,870-item multilingual dataset is publicly available.<\/p>\n","protected":false},"excerpt":{"rendered":"Ask an AI to estimate a startup&#8217;s odds of success. Now ask the same AI to estimate that&hellip;\n","protected":false},"author":2,"featured_media":124894,"comment_status":"","ping_status":"","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[8],"tags":[24,62804,53,3154,182,62806,11380,157,62805],"class_list":["post-124893","post","type-post","status-publish","format-standard","has-post-thumbnail","category-anthropic","tag-ai","tag-ai-probability-bias","tag-anthropic","tag-anthropic-claude","tag-claude","tag-claude-ai-alignment","tag-gpt","tag-openai","tag-optimismbench"],"_links":{"self":[{"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/posts\/124893","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/users\/2"}],"replies":[{"embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/comments?post=124893"}],"version-history":[{"count":0,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/posts\/124893\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/media\/124894"}],"wp:attachment":[{"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/media?parent=124893"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/categories?post=124893"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/tags?post=124893"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}