{"id":132558,"date":"2026-08-07T06:31:12","date_gmt":"2026-08-07T06:31:12","guid":{"rendered":"https:\/\/www.europesays.com\/ai\/132558\/"},"modified":"2026-08-07T06:31:12","modified_gmt":"2026-08-07T06:31:12","slug":"metas-muse-spark-1-2-matches-grok-4-5-in-overall-ai-performance-surging-in-just-four-months-biggo-finance","status":"publish","type":"post","link":"https:\/\/www.europesays.com\/ai\/132558\/","title":{"rendered":"Meta&#8217;s Muse Spark 1.2 Matches Grok 4.5 in Overall AI Performance, Surging in Just Four Months \u2014 BigGo Finance"},"content":{"rendered":"<p>Meta&#8217;s foundation model &#8220;Muse Spark 1.2,&#8221; announced on August 5, 2026, has achieved a high score on the &#8220;Artificial Analysis Intelligence Index&#8221;\u2014a benchmark that comprehensively evaluates AI capabilities in mathematics, science, coding, and reasoning\u2014surpassing SpaceXAI&#8217;s Grok 4.5. In just four months since the Muse series debuted, it has surged to tie for third place among U.S. companies, demonstrating explosive growth.<\/p>\n<p>According to Artificial Analysis, Muse Spark 1.2 scored &#8220;54&#8221; on its highest reasoning setting, &#8220;xhigh.&#8221; This represents a significant leap from Muse Spark 1.0&#8217;s &#8220;43&#8221; (released April 2026) and Muse Spark 1.1&#8217;s &#8220;51&#8221; (released July 2026) in a remarkably short period. The score is nearly on par with GPT-5.5 (xhigh) at &#8220;55&#8221; and Grok 4.5 (high) at &#8220;54,&#8221; tying with SpaceXAI for third place among U.S. AI companies.<\/p>\n<p>That said, a gap with the most advanced models persists. Muse Spark 1.2 trailed Claude Opus 5 (max) at &#8220;61,&#8221; Claude Fable 5 (fallback) at &#8220;60,&#8221; GPT-5.6 Sol (max) at &#8220;59,&#8221; and Kimi K3 (max) at &#8220;57.&#8221;<\/p>\n<p>Steady improvements were also observed in benchmarks measuring practical capabilities. On &#8220;GDPval-AA v2,&#8221; which assesses AI agents&#8217; real-world task execution abilities, the score rose from Muse Spark 1.1&#8217;s &#8220;1371&#8221; to &#8220;1631.&#8221; Only Claude Opus 5 (max) at &#8220;1852,&#8221; Claude Fable 5 (fallback) at &#8220;1743,&#8221; GPT-5.6 Sol (max) at &#8220;1730,&#8221; and Kimi K3 (max) at &#8220;1685&#8221; scored higher.<\/p>\n<p>On &#8220;Terminal-Bench v2.1,&#8221; which evaluates complex task execution in terminal environments, the score improved from 78% to 80%. On &#8220;\u03c4\u00b3-Banking,&#8221; measuring accuracy in financial customer support operations, it rose from 25% to 27%.<\/p>\n<p>However, scores declined on &#8220;SciCode,&#8221; which assesses numerical computation code generation for scientific research, dropping from 58.2% to 56.4%, and on &#8220;Humanity&#8217;s Last Exam,&#8221; which measures peak human-level expertise and reasoning ability, falling from 45.1% to 43.9%.<\/p>\n<p>On the &#8220;AA-Omniscience Index,&#8221; which gauges factual accuracy and hallucination suppression, the score rose from 18 to 22. While the hallucination rate decreased from 38% to 28%, the accuracy rate dipped from 41% to 38%.<\/p>\n<p>On the &#8220;Vals Index&#8221; conducted by Vals AI, Muse Spark 1.2 recorded the fifth-highest score at 71.88%. Notably, its cost per test was $0.69 (approximately \u00a5110), the cheapest among the top five AI models. This is roughly one-third the cost of Kimi K3 and less than one-tenth compared to Claude Fable 5, Opus 5, and GPT-5.6 Sol.<\/p>\n<p>Meta simultaneously unveiled &#8220;Muse Code (beta),&#8221; a terminal-based coding agent. Powered by Muse Spark 1.2, it outperformed OpenAI&#8217;s Codex and Google&#8217;s Antigravity on many coding tests but fell short of Anthropic&#8217;s Claude Opus 5 across all benchmarks.<\/p>\n<p>A standout feature of Muse Code is its crash-safe local event log. It records every model call, tool execution, approval, and edit, allowing users to resume from the point of interruption if a crash occurs. CEO Mark Zuckerberg explained, &#8220;When a job is large enough, it deploys to individual sub-agents working in parallel in isolated worktrees. It never touches the working copy.&#8221;<\/p>\n<p>The team behind &#8220;Cline,&#8221; an open-source AI coding agent, experimented by extracting instructions from Muse Spark 1.2&#8217;s system prompt and adding them to their own harness. They reported that token usage was reduced by 2.7 times, from 19.7 million to 7.2 million; completion time was halved from 49 minutes to 24 minutes; and cost was cut by 2.4 times, from $7.69 to $3.25.<\/p>\n<p>AI researcher Rihard Jarc commented, &#8220;With Muse Spark 1.2, Meta appears to have surpassed Google&#8217;s AI models in quality for many use cases. Considering the timeframe, this is truly astonishing.&#8221; He added, &#8220;Meta is already planning to release a model (codenamed: Watermelon) that outperforms Muse Spark, which should deliver Claude Fable-level performance.&#8221;<\/p>\n<p>Jarc also touched on speculation that DeepSeek is planning price hikes, noting, &#8220;Having sufficient computing resources to serve customers is critical. Meta is one of the few companies with computing power equal to or greater than Anthropic and OpenAI.&#8221;<\/p>\n<p>Meta has indicated to investors that it expects to invest between $125 billion (approximately \u00a519.8 trillion) and $145 billion (approximately \u00a523 trillion) in infrastructure such as chips and data centers this year, accelerating AI development backed by massive capital expenditure. While still trailing Opus 5 on benchmarks, the company is putting cost competitiveness front and center.<\/p>\n","protected":false},"excerpt":{"rendered":"Meta&#8217;s foundation model &#8220;Muse Spark 1.2,&#8221; announced on August 5, 2026, has achieved a high score on the&hellip;\n","protected":false},"author":2,"featured_media":132559,"comment_status":"","ping_status":"","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[10],"tags":[53,934,59173,132,6364,47109,1112,1122,65444,65746,157,19549,65977,2899],"class_list":["post-132558","post","type-post","status-publish","format-standard","has-post-thumbnail","category-xai","tag-anthropic","tag-artificial-analysis","tag-claude-opus-5","tag-google","tag-grok","tag-grok-4-5","tag-mark-zuckerberg","tag-meta","tag-muse-code","tag-muse-spark-1-2","tag-openai","tag-spacexai","tag-vals-ai","tag-xai"],"_links":{"self":[{"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/posts\/132558","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/users\/2"}],"replies":[{"embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/comments?post=132558"}],"version-history":[{"count":0,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/posts\/132558\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/media\/132559"}],"wp:attachment":[{"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/media?parent=132558"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/categories?post=132558"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/tags?post=132558"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}