{"id":118624,"date":"2026-07-25T10:14:13","date_gmt":"2026-07-25T10:14:13","guid":{"rendered":"https:\/\/www.europesays.com\/ai\/118624\/"},"modified":"2026-07-25T10:14:13","modified_gmt":"2026-07-25T10:14:13","slug":"anthropics-claude-opus-5-costs-well-below-fable-5-while-matching-or-beating-it-across-most-benchmarks","status":"publish","type":"post","link":"https:\/\/www.europesays.com\/ai\/118624\/","title":{"rendered":"Anthropic&#8217;s Claude Opus 5 costs well below Fable 5 while matching or beating it across most benchmarks"},"content":{"rendered":"<p>Anthropic&#8217;s Claude Opus 5 is the most capable AI model available today, according to several benchmarks, outperforming Fable 5 while costing less.<\/p>\n<p>Opus 5 scored 61 on the\u00a0<a href=\"https:\/\/artificialanalysis.ai\/\" target=\"_blank\" rel=\"noopener noreferrer nofollow\">Artificial Analysis Intelligence Index<\/a>, which combines nine tests covering knowledge work, coding, scientific reasoning, and factual accuracy. That puts it just ahead of Claude Fable 5 (60), GPT-5.6 Sol (59), Kimi K3 (57), and Claude Opus 4.8 (56). Artificial Analysis worked with Anthropic to test the model before its public release.<\/p>\n<p>In coding, Claude Opus 5 at &#8220;xhigh&#8221; paired with Claude Code shares first place on the Artificial Analysis Coding Index, which measures how well AI models handle programming tasks on their own, including finding and fixing bugs. On Terminal-Bench v2.1, which tests agents as autonomous engineers in real terminal environments, Opus 5 scored 89 percent at &#8220;max,&#8221; matching the previous leader GPT-5.6 Sol.<\/p>\n<p><img fetchpriority=\"high\" decoding=\"async\" class=\"wp-image-38198 size-full\" src=\"https:\/\/www.europesays.com\/ai\/wp-content\/uploads\/2026\/07\/AA_opus_5-1.png\" alt=\"\" width=\"1435\" height=\"1081\"\/>Claude Opus 5 at &#8220;max&#8221; leads the Intelligence Index with 61 points. On the Coding Agent Index, Opus 5 with Claude Code and GPT-5.6 Sol with Codex share first place with 67 points. | Image: Artificial Analysis<\/p>\n<p>For scientific reasoning, Opus 5 scored 53 percent on Humanity&#8217;s Last Exam, a very difficult knowledge test covering many academic fields. That ties it with Fable 5. On CritPt, a physics benchmark from researchers at Argonne National Laboratory and UIUC, it again matches Fable 5 but trails GPT-5.6 Sol, GPT-5.5 Pro, and GPT-5.6 Terra.<\/p>\n<p>Factual accuracy remains a weak spot: On AA-Omniscience, which tests the accuracy of a model&#8217;s knowledge claims, Opus 5 improved by 7 points over Opus 4.8 but still trails Fable 5. Opus 5 also answers more often when it&#8217;s uncertain, pushing its hallucination rate up 14 points to 50 percent.<\/p>\n<p>Epoch AI confirms tight race among frontier models<\/p>\n<p><a href=\"https:\/\/epoch.ai\/benchmarks\/eci?subset-view=graph&amp;subset-tab=Software+engineering&amp;view=graph&amp;tab=release-date\" target=\"_blank\" rel=\"noopener noreferrer nofollow\">Epoch AI<\/a> has also tested Claude Opus 5. The research institute gave it an overall Epoch Capability Index score of 159, just below Fable 5 at 161. When looking solely at the software engineering benchmarks (SWE-ECI), though, Opus 5 ties with Fable 5 at 161, outperforming GPT-5.6 Terra and Claude Opus 4.8. GPT-5.6 Sol leads in both categories.<\/p>\n<p><img loading=\"lazy\" decoding=\"async\" class=\"wp-image-38199 size-full\" src=\"https:\/\/www.europesays.com\/ai\/wp-content\/uploads\/2026\/07\/epoch_ai_ECI.jpg\" alt=\"\" width=\"1026\" height=\"988\"\/>Image: Epoch AI<\/p>\n<p>Overall, this confirms that the race among frontier models is tight. No single model can pull away or claim a clear advantage. That lends weight to the argument that <a href=\"https:\/\/the-decoder.com\/microsoft-ceo-satya-nadella-says-ai-models-are-getting-commoditized\/\" rel=\"nofollow noopener\" target=\"_blank\">AI models will eventually become commoditized<\/a>.<\/p>\n<p>Lower reasoning tiers deliver better value and coding results<\/p>\n<p>The average Intelligence Index task costs $2.03 with Opus 5, less than Claude Fable 5 with fallback at $2.75. It costs more than Opus 4.8 at $1.80 and Sonnet 5 at $1.53. At the &#8220;high&#8221; and &#8220;xhigh&#8221; tiers, however, Opus 5 beats both Opus 4.8 and Sonnet 5 while costing less.<\/p>\n<p><a href=\"https:\/\/vals.ai\/\" target=\"_blank\" rel=\"noopener noreferrer nofollow\">Vals.ai<\/a>\u00a0tested Claude Opus 5 across all five reasoning tiers using Vibe Code Bench, a benchmark for programming tasks. Scores climb from 76.7 percent at &#8220;low&#8221; to 82 percent at &#8220;medium&#8221; and 89.8 percent at &#8220;high.&#8221; Performance dips at the two highest tiers, with &#8220;xhigh&#8221; scoring 88.3 percent and &#8220;max&#8221; scoring 88.4 percent despite much higher costs.<\/p>\n<p>Vals.ai found that the highest tiers tend to produce more complex solutions that contain errors more often. The &#8220;high&#8221; tier produces simpler solutions that meet the requirements more reliably.<\/p>\n<p><a href=\"https:\/\/www.europesays.com\/ai\/wp-content\/uploads\/2026\/07\/opus5_vals_reasoning.jpg\"><img loading=\"lazy\" decoding=\"async\" class=\"size-full wp-image-38205\" src=\"https:\/\/www.europesays.com\/ai\/wp-content\/uploads\/2026\/07\/opus5_vals_reasoning.jpg\" alt=\"\" width=\"1200\" height=\"803\"\/><\/a>The default &#8220;high&#8221; reasoning tier appears to deliver the best results for Opus 5, at least for coding tasks. | Image: Vals.ai<\/p>\n<p><a target=\"_blank\" rel=\"noopener nofollow\" href=\"https:\/\/x.com\/ValsAI\/status\/2080817015871451500\">Terminal-Bench 2.1<\/a> shows a similar pattern. The &#8220;high&#8221; tier beats &#8220;max&#8221; because the model spends more time on each attempt at the top tier, leaving fewer attempts within the time limit. That matches Anthropic&#8217;s guidance, with &#8220;high&#8221; set as the default tier in both the API and Claude Code.<\/p>\n<p>Token pricing stays at $5 per million input tokens and $25 per million output tokens. Cache writes cost $6.25 per million tokens with a five-minute lifetime, while cache hits run just $0.50 per million tokens.<\/p>\n<p>Opus 5 pulls ahead in knowledge work<\/p>\n<p>Opus 5 performs especially well on the\u00a0<a href=\"https:\/\/artificialanalysis.ai\/evaluations\/aa-briefcase\" target=\"_blank\" rel=\"noopener noreferrer nofollow\">AA-Briefcase<\/a>\u00a0benchmark, which measures how well AI models handle typical office tasks like writing research reports, building presentations, and analyzing spreadsheets based on thousands of input files. Performance is scored across correctness, analytical quality, and presentation quality, then rolled into an Elo rating similar to chess rankings.<\/p>\n<p>At max reasoning, Opus 5 reaches an Elo of 1720, a full 146 points ahead of Claude Fable 5 (1574). Its three highest tiers (max, xhigh, high) sweep the top three spots. Combined with Fable 5, Sonnet 5, and Opus 4.8, Anthropic now holds the vast majority of top-10 positions.<\/p>\n<p><img loading=\"lazy\" decoding=\"async\" class=\"size-full wp-image-59225\" src=\"https:\/\/www.europesays.com\/ai\/wp-content\/uploads\/2026\/07\/AA_opus_5-2.png\" alt=\"\" width=\"1385\" height=\"1157\"\/>Opus 5 at &#8220;max&#8221; reaches 1861 Elo on GDPval-AA v2, far above the human baseline of 1000. Its three highest reasoning tiers also take the top three spots on AA-Briefcase. | Image: Artificial Analysis<\/p>\n<p>Cost per task dropped 20 percent to $17.79, down from $22.30 for Fable 5. The &#8220;xhigh&#8221; variant costs $14.26, and &#8220;high&#8221; comes in at just $10.41, less than half of Fable 5. Both still beat Fable 5 in the Elo ranking. At medium performance, Opus 5 hits an Elo of 1470, just behind GPT-5.6 Sol (max, 1505). At low performance, it lands at 1223 Elo, slightly below GLM-5.2 (max, 1254).<\/p>\n<p><img loading=\"lazy\" decoding=\"async\" class=\"size-full wp-image-59226\" src=\"https:\/\/www.europesays.com\/ai\/wp-content\/uploads\/2026\/07\/AA_opus_5-3.png\" alt=\"\" width=\"1437\" height=\"604\"\/>On AA-Briefcase, Opus 5 costs $17.79 per task at &#8220;max&#8221; and $10.41 at &#8220;high,&#8221; compared with $22.30 for Fable 5. | Image: Artificial Analysis<br \/>\n<img loading=\"lazy\" decoding=\"async\" class=\"size-full wp-image-59227\" src=\"https:\/\/www.europesays.com\/ai\/wp-content\/uploads\/2026\/07\/opus_5_token.png\" alt=\"\" width=\"1152\" height=\"528\"\/>Claude Opus 5 posts the highest GDPval-AA v2 scores but uses far more tokens than most rivals. Lower reasoning tiers move it into a more efficient range. | Image: Artificial Analysis<\/p>\n<p>Opus 5&#8217;s biggest gains show up in analytical quality. At &#8220;max,&#8221; it reaches an Analytical Quality Elo of 2016, almost 300 points ahead of Fable 5. Its Rubric Pass Rate, which tracks how often the model meets predefined quality standards, sits at 58 percent (&#8220;max&#8221;), 57.2 percent (&#8220;xhigh&#8221;), and 56 percent (&#8220;high&#8221;). Presentation quality is a different story. Opus 5 scores a Presentation Elo of 1628, about 40 points behind GPT-5.6 Sol at &#8220;max&#8221; (1666).<\/p>\n<p>More performance takes more time: At &#8220;max,&#8221; Opus 5 needs over 36 minutes per task and averages 103 passes, about 50 percent longer than Opus 4.8 at 24 minutes and 55 passes.<\/p>\n<p>\t\t\t\tAI News Without the Hype \u2013 Curated by Humans<\/p>\n<p>\n\t\t\t\t\tSubscribe to THE DECODER for ad-free reading, a weekly AI newsletter, our exclusive &#8220;AI Radar&#8221; frontier report six times a year, full archive access, and access to our comment section.\t\t\t\t<\/p>\n<p>\t\t\t\t<a href=\"https:\/\/the-decoder.com\/subscription\/\" class=\"inline-block text-white bg-(--heise-primary) mt-3 hover:bg-blue-800 focus:ring-4 focus:outline-none focus:ring-blue-300 font-medium rounded-sm w-full sm:w-auto  pl-3 pr-3 py-2.5 text-center newsletter-submit-button hover:no-underline\" rel=\"nofollow noopener\" target=\"_blank\"><br \/>\n\t\t\t\t\tSubscribe now\t\t\t\t<\/a><\/p>\n","protected":false},"excerpt":{"rendered":"Anthropic&#8217;s Claude Opus 5 is the most capable AI model available today, according to several benchmarks, outperforming Fable&hellip;\n","protected":false},"author":2,"featured_media":118625,"comment_status":"","ping_status":"","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[8],"tags":[53,3154,182,59173,59860],"class_list":["post-118624","post","type-post","status-publish","format-standard","has-post-thumbnail","category-anthropic","tag-anthropic","tag-anthropic-claude","tag-claude","tag-claude-opus-5","tag-opus-5"],"_links":{"self":[{"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/posts\/118624","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/users\/2"}],"replies":[{"embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/comments?post=118624"}],"version-history":[{"count":0,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/posts\/118624\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/media\/118625"}],"wp:attachment":[{"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/media?parent=118624"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/categories?post=118624"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/tags?post=118624"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}