Meta’s foundation model “Muse Spark 1.2,” announced on August 5, 2026, has achieved a high score on the “Artificial Analysis Intelligence Index”—a benchmark that comprehensively evaluates AI capabilities in mathematics, science, coding, and reasoning—surpassing SpaceXAI’s Grok 4.5. In just four months since the Muse series debuted, it has surged to tie for third place among U.S. companies, demonstrating explosive growth.
According to Artificial Analysis, Muse Spark 1.2 scored “54” on its highest reasoning setting, “xhigh.” This represents a significant leap from Muse Spark 1.0’s “43” (released April 2026) and Muse Spark 1.1’s “51” (released July 2026) in a remarkably short period. The score is nearly on par with GPT-5.5 (xhigh) at “55” and Grok 4.5 (high) at “54,” tying with SpaceXAI for third place among U.S. AI companies.
That said, a gap with the most advanced models persists. Muse Spark 1.2 trailed Claude Opus 5 (max) at “61,” Claude Fable 5 (fallback) at “60,” GPT-5.6 Sol (max) at “59,” and Kimi K3 (max) at “57.”
Steady improvements were also observed in benchmarks measuring practical capabilities. On “GDPval-AA v2,” which assesses AI agents’ real-world task execution abilities, the score rose from Muse Spark 1.1’s “1371” to “1631.” Only Claude Opus 5 (max) at “1852,” Claude Fable 5 (fallback) at “1743,” GPT-5.6 Sol (max) at “1730,” and Kimi K3 (max) at “1685” scored higher.
On “Terminal-Bench v2.1,” which evaluates complex task execution in terminal environments, the score improved from 78% to 80%. On “τ³-Banking,” measuring accuracy in financial customer support operations, it rose from 25% to 27%.
However, scores declined on “SciCode,” which assesses numerical computation code generation for scientific research, dropping from 58.2% to 56.4%, and on “Humanity’s Last Exam,” which measures peak human-level expertise and reasoning ability, falling from 45.1% to 43.9%.
On the “AA-Omniscience Index,” which gauges factual accuracy and hallucination suppression, the score rose from 18 to 22. While the hallucination rate decreased from 38% to 28%, the accuracy rate dipped from 41% to 38%.
On the “Vals Index” conducted by Vals AI, Muse Spark 1.2 recorded the fifth-highest score at 71.88%. Notably, its cost per test was $0.69 (approximately ¥110), the cheapest among the top five AI models. This is roughly one-third the cost of Kimi K3 and less than one-tenth compared to Claude Fable 5, Opus 5, and GPT-5.6 Sol.
Meta simultaneously unveiled “Muse Code (beta),” a terminal-based coding agent. Powered by Muse Spark 1.2, it outperformed OpenAI’s Codex and Google’s Antigravity on many coding tests but fell short of Anthropic’s Claude Opus 5 across all benchmarks.
A standout feature of Muse Code is its crash-safe local event log. It records every model call, tool execution, approval, and edit, allowing users to resume from the point of interruption if a crash occurs. CEO Mark Zuckerberg explained, “When a job is large enough, it deploys to individual sub-agents working in parallel in isolated worktrees. It never touches the working copy.”
The team behind “Cline,” an open-source AI coding agent, experimented by extracting instructions from Muse Spark 1.2’s system prompt and adding them to their own harness. They reported that token usage was reduced by 2.7 times, from 19.7 million to 7.2 million; completion time was halved from 49 minutes to 24 minutes; and cost was cut by 2.4 times, from $7.69 to $3.25.
AI researcher Rihard Jarc commented, “With Muse Spark 1.2, Meta appears to have surpassed Google’s AI models in quality for many use cases. Considering the timeframe, this is truly astonishing.” He added, “Meta is already planning to release a model (codenamed: Watermelon) that outperforms Muse Spark, which should deliver Claude Fable-level performance.”
Jarc also touched on speculation that DeepSeek is planning price hikes, noting, “Having sufficient computing resources to serve customers is critical. Meta is one of the few companies with computing power equal to or greater than Anthropic and OpenAI.”
Meta has indicated to investors that it expects to invest between $125 billion (approximately ¥19.8 trillion) and $145 billion (approximately ¥23 trillion) in infrastructure such as chips and data centers this year, accelerating AI development backed by massive capital expenditure. While still trailing Opus 5 on benchmarks, the company is putting cost competitiveness front and center.