{"id":118416,"date":"2026-07-25T01:44:07","date_gmt":"2026-07-25T01:44:07","guid":{"rendered":"https:\/\/www.europesays.com\/ai\/118416\/"},"modified":"2026-07-25T01:44:07","modified_gmt":"2026-07-25T01:44:07","slug":"claude-opus-5-outscores-fable-5-on-8-of-13-benchmarks-at-half-the-token-price","status":"publish","type":"post","link":"https:\/\/www.europesays.com\/ai\/118416\/","title":{"rendered":"Claude Opus 5 outscores Fable 5 on 8 of 13 benchmarks at half the token price"},"content":{"rendered":"<p>Anthropic\u2019s Claude Fable 5 is arguably the most influential LLM to launch in recent memory, marking the biggest leap in capabilities since Opus 4.5 debuted in November 2025. The publicly available counterpart to Claude Mythos Preview, which had a sizable jump in capabilities, most notably, in cybersecurity, Claude Fable featured stronger agentic performance and markedly better performance in coding tasks.<\/p>\n<p>But Fable had a difficult forty-five days. After initially launching on June 9, it went dark internationally <a href=\"https:\/\/www.anthropic.com\/news\/fable-mythos-access\" rel=\"nofollow noopener\" target=\"_blank\">three days later under Commerce Department export controls<\/a>. The model returned July 1, and by July 22 was <a href=\"https:\/\/techcrunch.com\/2026\/07\/22\/treasury-threatens-sanctions-after-white-house-claims-moonshot-distilled-anthropics-fable\/\" rel=\"nofollow noopener\" target=\"_blank\">the subject of a White House accusation<\/a> that a Chinese lab had covertly distilled it. Not everyone says it is that simple.\u00a0Braden Hancock of the Laude Institute, for instance, <a href=\"https:\/\/techcrunch.com\/2026\/07\/23\/experts-say-exploiting-anthropics-fable-isnt-how-kimi-k3-got-so-good\/\" rel=\"nofollow noopener\" target=\"_blank\">told TechCrunch<\/a> that a model this strong could not come from straight distillation with Fable only public since July 1.<\/p>\n<p>In any case, on Friday, Anthropic released a model that beats Fable on most of the company\u2019s own benchmarks for half the price.<\/p>\n<p><img loading=\"lazy\" decoding=\"async\" class=\"size-large wp-image-82442 aligncenter\" src=\"https:\/\/www.europesays.com\/ai\/wp-content\/uploads\/2026\/07\/opus5-vs-fable5-head-to-head-1024x772.png\" alt=\"\" width=\"740\" height=\"558\"  \/><\/p>\n<p>In terms of pricing, Opus 5 lands at $5 per million input tokens and $25 per million output tokens. That is identical to Opus 4.8 and half what Anthropic charges for Fable 5. GPT-5.6 Sol is in the same ballpark at $5 and $30. Kimi K3, the most expensive open-weight model (<a href=\"https:\/\/www.techi.com\/kimi-k3-open-weights-inference-economics\/\" rel=\"nofollow noopener\" target=\"_blank\">weights to drop on July 27<\/a>), undercuts both at $3 and $15, with cache hits billed at $0.30.<\/p>\n<p><img loading=\"lazy\" decoding=\"async\" class=\"size-large wp-image-82445 aligncenter\" src=\"https:\/\/www.europesays.com\/ai\/wp-content\/uploads\/2026\/07\/frontier-price-vs-index-1024x683.png\" alt=\"\" width=\"740\" height=\"494\"  \/><\/p>\n<p>In a private head-to-head testing on a narrow language-engineering task, GPT-5.6 Sol at medium effort matched or beat Claude Opus 5 at every effort tier we tested. (Caveat: This test didn\u2019t test Opus 5\u2019s top two effort tiers: xhigh and max.) The task involved mapping audio files to corresponding text. Sol mapped the recordings to the correct printed section 100% of the time on both units. Opus 5 managed 100% on one and 88.4% on the other, where a single misplaced section boundary swept 14 consecutive recordings into the neighboring drill. Transcription accuracy favored Sol as well, 96.1% to 90.7% on the harder unit. The prompt was authored and iterated against GPT-5.6, then replayed to Opus 5 verbatim.<\/p>\n<p>Opus 5 bills less per output token than Sol under batch pricing, $12.50 per million against $15, yet cost 21% more per unit at its cheapest setting and 80% more at high effort, because it generated 1.5 to 2.5 times as many output tokens. Adaptive thinking is on by default in Opus 5. Per-token price turned out to be a poor predictor of per-task cost.<\/p>\n<p>Anthropic\u2019s own migration guide in the <a href=\"https:\/\/www-cdn.anthropic.com\/b514064af1408018e64b1ad24e7d5e75850b4ffd\/Claude%20Opus%205%20System%20Card.pdf\" rel=\"nofollow noopener\" target=\"_blank\">Opus 5 system card<\/a> warns that max effort can produce diminishing returns and overthinking on simpler tasks. On FrontierCode, scores fell above high effort because the model made unnecessary refactors and other out-of-scope edits. A brief instruction to stay within the requested scope recovered performance on most tasks. External pilot users likewise reported cases where higher effort made the model perform worse.<\/p>\n<p>More thinking sometimes also means more opportunities for hallucinations.\u00a0On the closed-book <a href=\"https:\/\/artificialanalysis.ai\/evaluations\/omniscience\" rel=\"nofollow noopener\" target=\"_blank\">AA-Omniscience benchmark<\/a>, Opus 5 was 11% more accurate than Opus 4.8, while its hallucination rate was 6% higher.<\/p>\n","protected":false},"excerpt":{"rendered":"Anthropic\u2019s Claude Fable 5 is arguably the most influential LLM to launch in recent memory, marking the biggest&hellip;\n","protected":false},"author":2,"featured_media":118417,"comment_status":"","ping_status":"","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[8],"tags":[179,2403,2720,59946,7759,53,3154,182,38051,38060,59173,59947,9878,46529,56012,1642,59948,1738,720,59949],"class_list":["post-118416","post","type-post","status-publish","format-standard","has-post-thumbnail","category-anthropic","tag-agentic-ai","tag-ai-benchmarks","tag-ai-hallucinations","tag-ai-model-safety","tag-ai-reasoning","tag-anthropic","tag-anthropic-claude","tag-claude","tag-claude-fable-5","tag-claude-mythos-5","tag-claude-opus-5","tag-coding-models","tag-frontier-ai-models","tag-gpt-5-6-sol","tag-kimi-k3","tag-large-language-models","tag-llm-pricing","tag-multi-agent-systems","tag-prompt-injection","tag-test-time-compute"],"_links":{"self":[{"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/posts\/118416","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/users\/2"}],"replies":[{"embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/comments?post=118416"}],"version-history":[{"count":0,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/posts\/118416\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/media\/118417"}],"wp:attachment":[{"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/media?parent=118416"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/categories?post=118416"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/tags?post=118416"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}