{"id":149572,"date":"2026-08-24T16:06:23","date_gmt":"2026-08-24T16:06:23","guid":{"rendered":"https:\/\/www.europesays.com\/ai\/149572\/"},"modified":"2026-08-24T16:06:23","modified_gmt":"2026-08-24T16:06:23","slug":"nvidia-avo-pushes-claude-opus-5-to-a-perfect-arc-agi-3-benchmark-score","status":"publish","type":"post","link":"https:\/\/www.europesays.com\/ai\/149572\/","title":{"rendered":"NVIDIA AVO Pushes Claude Opus 5 To A Perfect ARC-AGI-3 Benchmark Score"},"content":{"rendered":"<p><img decoding=\"async\" class=\" top-image\" src=\"https:\/\/www.europesays.com\/ai\/wp-content\/uploads\/2026\/08\/1787587583_352_0x0.jpg\" alt=\"FRANCE-TECHNOLOGY-INDUSTRY-SCIENCE-FAIR-VIVATECH\" data-height=\"1652\" data-width=\"2478\" fetchpriority=\"high\" style=\"position:absolute;top:0\"\/><\/p>\n<p>NVIDIA components are displayed at the GTC Paris NVIDIA at the VivaTech technology startups and innovation fair at the Paris Expo Porte de Versailles, in Paris on June 12, 2025. The VivaTech fair opened in Paris on June 11, 2025 in the presence of the French Minister for Digital Technologies, before welcoming a number of tech stars and the French President against a backdrop of trade tensions between Europe and the United States. (Photo by Thomas SAMSON \/ AFP via Getty Images)<\/p>\n<p>AFP via Getty Images<\/p>\n<p>Nvidia took a frontier AI model that completes about 30 percent of one of the field\u2019s hardest benchmarks and got a perfect score out of it.<\/p>\n<p>The model was Anthropic\u2019s Claude Opus 5, and nothing inside it changed: no retraining, no fine-tuning, not one adjusted weight. Everything that improved sat outside the model, in a software system Nvidia calls AVO that manages the model\u2019s memory, plans its next moves, and watches for mistakes. Better software pulled three times more performance out of a model that already exists.<\/p>\n<p>What AVO Actually Is<\/p>\n<p>The software layer between a model and a task is called a harness, and Nvidia\u2019s researchers describe its job in one clean passage:<\/p>\n<p>A frontier language model is only one component of an AI agent. The surrounding agent system\u2014often called a harness\u2014determines how the model receives context, uses tools, maintains state, responds to feedback, recovers from failure, and sustains progress over long-running tasks.<\/p>\n<p>AVO, short for Agentic Variation Operators, is Nvidia\u2019s harness, and it has four working parts:<\/p>\n<p>A persistent memory carries forward everything the agent has tried, learned, and reasoned through, so it picks up where it left off instead of rebuilding its understanding from scratch.A supervisor watches the main agent the way a manager watches a new hire, stepping in when progress stalls and redirecting it toward a different strategy.An iterative loop drives the work: inspect the situation, plan, act, evaluate the result, repeat.Swappable tools adapt the system to whatever job it is given.<\/p>\n<p>The swappable tools are what make AVO general rather than a one-benchmark trick. Nvidia first built the system to optimize its own GPU code, then replaced the code tools with game controls and pointed it at a completely different job, where it performed at the frontier again.<\/p>\n<p>The model supplies the intelligence. The harness keeps that intelligence pointed at the job long enough to finish it.<\/p>\n<p>The benchmark comes from the ARC Prize team, and it is built to resist what AI models are usually good at. ARC-AGI-3 drops an agent into game-like worlds with no instructions, no explicit rules, and no stated goal.<\/p>\n<p>The agent has to figure out what each game even wants through trial and error, the way a person handed an unlabeled puzzle would. Memorized knowledge does not help, because the agent has never seen these games before. What gets measured is learning on the fly, the skill long business tasks actually require.<\/p>\n<p>The public set holds 25 environments of six to ten levels each, 183 levels in all, and scoring combines completion with action efficiency, benchmarked against humans playing the same games for the first time. AVO, running Claude Opus 5 underneath, completed every level in 6,624 actions, about 12 percent fewer than the previous most efficient system. <a href=\"https:\/\/thenewstack.io\/nvidia-avo-arcagi3-benchmark\/\" target=\"_blank\" rel=\"nofollow noopener noreferrer\" data-ga-track=\"ExternalLink:https:\/\/thenewstack.io\/nvidia-avo-arcagi3-benchmark\/\" aria-label=\"The bare model completes about 30 percent\">The bare model completes about 30 percent<\/a>. These are the benchmark\u2019s public environments; the private competition sets remain the harder proving ground.<\/p>\n<p>The Pilot Problem<\/p>\n<p>Enterprises do not buy benchmark scores. They buy work that gets finished, and that is where AI adoption has been stuck.<\/p>\n<p>MIT\u2019s NANDA initiative found that only about 5 percent of the enterprise AI initiatives it examined produced measurable business value, despite an estimated $30 billion to $40 billion in spending. The failures were rarely about model quality. Companies could not get AI to fit into existing work, remember what it had learned, or improve with feedback. An AI system that is expensive to run and unreliable is a business liability, not a productivity tool.<\/p>\n<p>A harness attacks that failure mode directly. Memory means the agent stops repeating yesterday\u2019s mistakes. Supervision means it stops burning money on dead ends. Fewer wasted actions lower the cost of every completed task while raising the quality of the result.<\/p>\n<p>Companies do not need to wait for a new generation of models to expand what AI does for them. The models already in production carry far more capability than current software extracts, and drawing it out costs a fraction of what training a new frontier model does.<\/p>\n<p>None of this shrinks the compute story. An agent that works a task for hours consumes far more inference than a chatbot answering a question, so dependable agents grow the compute bill. Each big drop in the cost of computing has expanded how much of it gets used. Reliability works the same way.<\/p>\n<p>Where The Value Lands<\/p>\n<p>The investor question is who profits when AI moves from impressive demos to dependable work. Companies like Microsoft, ServiceNow, Salesforce, and Palantir already sit inside the workflows enterprises want agents to run, which puts them exactly where the harness meets real work. The harness needs the context those platforms own: the documents, the tickets, the customer records, the code.<\/p>\n<p>Then there is the company that ran the experiment. Nvidia built AVO in-house, an agent system it also uses to optimize its own GPU code. The same week the benchmark result landed, Nvidia reportedly agreed to pay $6 billion to license the model-building software of the coding startup Poolside. The company that sells the chips is assembling the software layers above them, one deliberate move at a time.<\/p>\n<p>Enterprises stalled on AI because the jobs rarely got finished. The software that finishes them has started to arrive, and the companies supplying it are the ones positioned to collect.<\/p>\n","protected":false},"excerpt":{"rendered":"NVIDIA components are displayed at the GTC Paris NVIDIA at the VivaTech technology startups and innovation fair at&hellip;\n","protected":false},"author":2,"featured_media":149573,"comment_status":"","ping_status":"","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[8],"tags":[405,407,1276,53,3154,60278,73204,182,59173,2407,64139,73203],"class_list":["post-149572","post","type-post","status-publish","format-standard","has-post-thumbnail","category-anthropic","tag-ai-agents","tag-ai-infrastructure","tag-ai-models","tag-anthropic","tag-anthropic-claude","tag-arc-agi-3","tag-avo-harness","tag-claude","tag-claude-opus-5","tag-frontier-ai","tag-llm-benchmark","tag-nvidia-avo"],"_links":{"self":[{"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/posts\/149572","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/users\/2"}],"replies":[{"embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/comments?post=149572"}],"version-history":[{"count":0,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/posts\/149572\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/media\/149573"}],"wp:attachment":[{"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/media?parent=149572"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/categories?post=149572"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/tags?post=149572"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}