Anthropic has rolled out its latest mid-tier artificial intelligence model, Claude Opus 5, in a move that brings near-flagship capabilities to a much broader user base at a significantly lower price point. The San Francisco-based startup launched the model on Thursday, positioning it as the new workhorse for enterprise coding, knowledge work, and scientific research.
According to Anthropic, Opus 5 delivers performance that closely rivals—and in several benchmarks even surpasses—its most powerful commercially available model, Fable 5, while costing customers exactly half the price. The API pricing remains unchanged from its predecessor, Opus 4.8, at $5 per million input tokens and $25 per million output tokens. In contrast, Fable 5 costs $10 per million input tokens and $50 per million output tokens.
Anthropic product leader Dianne Penn told Reuters that the release reflects the company’s rapid pace of development. “We’re building and continue to consistently deliver frontier intelligence and bring that as accessibly as possible with every model generation,” Penn said.
The new model becomes the default for Claude Max subscribers and the most capable option available on the Claude Pro plan. Penn noted that while Opus 5 approaches Fable 5 on many tasks, users running “days-long, very autonomous projects” should still opt for the higher-end model.
Coding and Knowledge Work Benchmarks
Anthropic’s internal testing shows substantial gains for Opus 5 over its predecessor and, in some cases, over Fable 5 itself. On the agentic terminal coding evaluation Frontier Bench v0.1, Opus 5 scored 43.3%, compared to 33.7% for Fable 5, 21.1% for Opus 4.8, and 34.4% for OpenAI’s GPT-5.6 Sol. On the knowledge work benchmark GDPVal-AA v2, the model achieved 1,861 points, the highest among all models tested.
The model also demonstrated significant progress in novel problem-solving. On the ARC-AGI 3 benchmark, Opus 5 scored 30.2%, dramatically ahead of GPT-5.6 Sol’s 7.8% and Opus 4.8’s 1.5%. In computer-use evaluation OSWorld 2.0 and enterprise automation benchmark Automation Bench, Opus 5 scored 70.6% and 26.0% respectively, leading the field.
BenchmarkOpus 5Fable 5Opus 4.8GPT-5.6 SolFrontier Bench v0.1 (Coding)43.3%33.7%21.1%34.4%GDPVal-AA v2 (Knowledge)1,861———ARC-AGI 3 (Reasoning)30.2%—1.5%7.8%OSS-Fuzz (Vulnerability Detection)79.4%———
註:Anthropic 並未公開 Fable 5 和 Opus 4.8 在 GDPVal-AA v2 的具體分數;Mythos 5 的 OSS-Fuzz 分數為 80.0%。
Scientific Research and Cybersecurity
In the life sciences, Opus 5 outperformed Opus 4.8 across structural biology, organic chemistry, and bioinformatics. An internal evaluation on inferring molecular structures from spectroscopic data showed a 10.2 percentage point improvement, while predicting the functional impact of protein sequence variations saw a 7.7 point gain.
Cybersecurity is an area where Anthropic has navigated a delicate balance. Opus 5 achieved a 79.4% score on the OSS-Fuzz benchmark for finding vulnerabilities in open-source code, nearly matching the 80.0% of the restricted Mythos 5 model. However, it succeeded in exploiting those vulnerabilities in only 4 out of 14 attempts, compared to 13 for Mythos 5. Anthropic stated it did not specifically train Opus 5 on offensive cyber operations and restricts requests for penetration testing and exploit generation.
Because Opus 5 is less capable of weaponizing code flaws, its safeguards are proportionally less restrictive than those on Fable 5. Anthropic expects that safety classifiers will intervene roughly 85% less often for Opus 5 than for Fable 5, a welcome change for users who found Fable’s guardrails too aggressive. The model is also not subject to the 30-day data retention policy that applies to Fable and Mythos 5, addressing privacy concerns raised by some enterprise customers.
Safety, Alignment, and New Features
Anthropic described Opus 5 as its “most aligned model to date,” with the lowest rates of deceptive behavior and the least susceptibility to being tricked into misuse among its current models. Automated behavioral testing yielded a score of 2.30, compared to 2.85 for Opus 4.8, 2.81 for Fable 5, and 3.35 for Sonnet 5, with lower scores indicating better safety.
Alongside the model, Anthropic introduced several beta features. A “Fast Mode” runs approximately 2.5 times quicker than the default speed at double the standard price. An “Effort” setting allows users to instruct the model to prioritize thoroughness or conserve tokens. The company also launched Automatic Fallbacks, an API feature that reroutes requests flagged by safety filters to a less capable model instead of blocking them outright, and mid-conversation tool-switching capabilities.
Market Context and Competition
The launch comes amid intensifying price competition in the AI sector. China-based Moonshot recently released its open-weight Kimi K3 model, priced at $15 per million output tokens, which has drawn attention from both users and U.S. regulators. Asked about Kimi K3, Penn told Reuters that it “remains to be seen” how open-weight models perform on complex, real-world projects.
According to data from Artificial Analysis, the weighted average cost per Intelligence Index task is $2.03 for Opus 5 at maximum settings, compared to $2.75 for Fable 5, $1.04 for GPT-5.6 Sol, and $0.95 for Kimi K3. Opus 5 leads the Artificial Analysis Intelligence Index with a score of 61, one point ahead of Fable 5.
By maintaining the price of the previous-generation Opus while delivering performance that often exceeds Fable 5 on coding and knowledge work, Anthropic is directly targeting enterprises that prioritize cost efficiency and practical utility over raw, unrestricted power. The strategy signals a new phase in the AI arms race, where the ability to deploy near-frontier intelligence economically is becoming as critical as achieving state-of-the-art benchmarks.