{"id":149546,"date":"2026-08-24T15:37:09","date_gmt":"2026-08-24T15:37:09","guid":{"rendered":"https:\/\/www.europesays.com\/ai\/149546\/"},"modified":"2026-08-24T15:37:09","modified_gmt":"2026-08-24T15:37:09","slug":"nvidia-vera-rubin-nvl72-sets-a-new-efficiency-standard-for-ai-agents","status":"publish","type":"post","link":"https:\/\/www.europesays.com\/ai\/149546\/","title":{"rendered":"NVIDIA Vera Rubin NVL72 Sets a New Efficiency Standard for AI Agents"},"content":{"rendered":"<p>According to <a target=\"_blank\" href=\"https:\/\/openrouter.ai\/blog\/insights\/deepseek-v4-adoption\/\" rel=\"nofollow noopener\">OpenRouter data<\/a>, agentic AI workloads consume 15x more tokens than a simple chat request. Why?\u00a0<\/p>\n<p>Consider what happens when an AI agent researches a company for an investment decision. The agent queries financial databases, searches news and filings, invokes a sub-agent to run peer comparisons and model valuations, then synthesizes everything into a recommendation. Agents and sub-agents keep reasoning until the task is done, driving increased token demand. With every step, the accumulated tokens become the input to the next, making long-context handling central to agentic AI performance.<\/p>\n<p><img loading=\"lazy\" decoding=\"async\" class=\"alignnone wp-image-97855 size-full\" src=\"https:\/\/www.europesays.com\/ai\/wp-content\/uploads\/2026\/08\/end-to-end-corp-blog-diagram-agentx-activation-1280x680-5565900.gif\" alt=\"\" width=\"1280\" height=\"680\"\/><\/p>\n<p>The same pattern plays out across every agentic use case, from software development to customer service to deep research.\u00a0<\/p>\n<p>As agentic AI moves into production across industries, the infrastructure running it needs to meet that token demand efficiently.\u00a0<\/p>\n<p>New measured performance data shows NVIDIA Vera Rubin NVL72 systems deliver up to 30x higher throughput per megawatt than NVIDIA GB300 NVL72 on agentic workloads. NVIDIA measured this inference throughput data using the SemiAnalysis AgentX workload, consisting of recorded real-world agentic coding sessions, with actual context growth, tool calls and sub-agent spawning preserved. For power-constrained AI factories, that translates directly into 30x more agentic work for the same energy footprint.<\/p>\n<p>These early results for Vera Rubin NVL72 demonstrate NVIDIA\u2019s accelerated pace of innovation. With continuous software optimizations, performance across both Vera Rubin NVL72 and GB300 NVL72 will continue to improve.<\/p>\n<p>Vera Rubin NVL72: 30x Higher Throughput per Megawatt and 35x Lower Token Cost<\/p>\n<p>Agentic workloads look fundamentally different from chat or document summarization, where input and output sequences typically range from 1K to 8K tokens. In agentic sessions, context accumulates across steps and can reach hundreds of thousands of input tokens, with wide variability in both input and output lengths across requests. Performance measurement must evolve to capture the full agent workflow rather than a single inference request.<\/p>\n<p><img loading=\"lazy\" decoding=\"async\" class=\"alignnone size-full wp-image-97856\" src=\"https:\/\/www.europesays.com\/ai\/wp-content\/uploads\/2026\/08\/InferenceNeedsAgenticBenchmarks-scaled.png\" alt=\"\" width=\"2048\" height=\"1152\"  \/><\/p>\n<p>The results below reflect performance measured on real-world agentic coding trajectories.<\/p>\n<p>In <a target=\"_blank\" href=\"https:\/\/newsletter.semianalysis.com\/p\/agentx-inferencexv3-does-cuda-moat\" rel=\"nofollow noopener\">SemiAnalysis AgentX<\/a>, the NVIDIA Blackwell platform delivers leading performance across multiple agentic models including Kimi K3, MiniMax M3, GLM5.3, Qwen3.5 and DeepSeek V4 Pro.\u00a0<\/p>\n<p>For example, <a target=\"_blank\" href=\"https:\/\/developer.nvidia.com\/blog\/nvidia-vera-rubin-and-blackwell-set-a-new-standard-for-agentic-ai-performance-per-watt\/?ncid=so-inst-151306\" rel=\"nofollow noopener\">GB300 NVL72 delivers up to 15x<\/a> better throughput per megawatt than the NVIDIA Hopper architecture on the DeepSeek V4 Pro model, giving customers a high-performance foundation to run agentic workloads. This leap reflects the advantage of a larger scale-up GPU domain and codesigned software in delivering significantly better inference efficiency.<\/p>\n<p>Vera Rubin extends that advantage, lifting the performance across the entire Pareto curve, to deliver as much as 30x higher throughput per megawatt than GB300 NVL72 on the DeepSeek V4 Pro model. These early results, measured using the SemiAnalysis AgentX workload and currently pending SemiAnalysis review, don\u2019t yet reflect Vera CPU performance for tool calling.\u00a0<\/p>\n<p><img loading=\"lazy\" decoding=\"async\" class=\"alignnone size-full wp-image-97890\" src=\"https:\/\/www.europesays.com\/ai\/wp-content\/uploads\/2026\/08\/end-to-end-social-chart1-agentx-activation-s1-1920x1080-1.jpg\" alt=\"\" width=\"1920\" height=\"1080\"  \/><\/p>\n<p>NVIDIA DSX MaxLPS technologies manages power across the GPU, rack and workload levels to provision up to 40% more GPUs within the same megawatt budget, pushing throughput per megawatt further at AI factory scale.\u00a0<\/p>\n<p>Throughput per megawatt also directly impacts the cost of every token produced. At up to 35x lower cost per million tokens than GB300 NVL72, Vera Rubin NVL72 can run agents continuously, at scale, across the full breadth of customers\u2019 workloads.\u00a0<\/p>\n<p><img loading=\"lazy\" decoding=\"async\" class=\"alignnone size-full wp-image-97891\" src=\"https:\/\/www.europesays.com\/ai\/wp-content\/uploads\/2026\/08\/end-to-end-social-chart1-agentx-activation-s2-1920x1080-1.jpg\" alt=\"\" width=\"1920\" height=\"1080\"  \/><\/p>\n<p>For power-constrained AI factories, throughput per megawatt determines <a target=\"_blank\" href=\"https:\/\/www.cio.com\/article\/4209776\/5-critical-questions-that-define-ai-factory-economics.html\" rel=\"nofollow noopener\">AI factory revenue<\/a> and cost per million tokens determines the profit margin on that revenue.<\/p>\n<p>Extreme Codesign for Agentic Scale<\/p>\n<p>Modern inference optimization spans a range of techniques that are especially critical for agentic AI. NVIDIA Vera Rubin NVL72 enables all of these and more through extreme codesign across every layer of the platform to deliver multifold performance gains.<\/p>\n<p>Disaggregated serving separates context processing (prefill) from response generation (decode) so each scales independently.<br \/>\nRate matching synchronizes the speeds at which prefill GPUs and decode GPUs produce tokens to maximize efficiency.<br \/>\nLarge-scale expert parallelism distributes expert sub-networks in <a href=\"https:\/\/blogs.nvidia.com\/blog\/mixture-of-experts-frontier-models\/\" rel=\"nofollow noopener\" target=\"_blank\">mixture-of-experts models<\/a> across the scale-up GPU domain.\u00a0<br \/>\nDistributed KV-caching extends memory across the scale-up GPU domain, while KV-cache offloading tiers less-active context to host and storage, keeping previously processed context accessible without recomputation.<br \/>\nKV-aware routing directs incoming requests to the GPUs that already hold the relevant cached context, reducing redundant computation across long sessions.\u00a0<br \/>\nFused CUDA kernels like MegaMoE combine many computation and inter-GPU communication operations into a single execution pass, keeping GPUs active rather than waiting for data.<\/p>\n<p>NVIDIA Rubin GPUs\u2019 enhanced fifth-generation <a target=\"_blank\" href=\"https:\/\/www.nvidia.com\/en-us\/data-center\/tensor-cores\/\" rel=\"nofollow noopener\">Tensor Cores<\/a> and the third-generation <a target=\"_blank\" href=\"https:\/\/www.nvidia.com\/en-us\/data-center\/tensor-cores\/\" rel=\"nofollow noopener\">Transformer Engine<\/a> accelerate both prefill and decode stages of inference. NVFP4 quantization compresses model weights to 4-bit precision, reducing memory footprint and increasing throughput without sacrificing output quality.\u00a0<\/p>\n<p>The NVL72 scale-up domain, a defining architecture across Vera Rubin and Grace Blackwell, enables the high-bandwidth and low-latency inter-GPU communication essential for techniques such as large-scale expert parallelism and distributed KV-caching. Purpose-built to power this scale-up domain, NVIDIA NVLink interconnect technology and NVLink Switches, now in their sixth generation, deliver 10x higher packet rates and 3x lower latency than off-the-shelf Ethernet alternatives.<\/p>\n<p>Spanning optimized CUDA kernels, inference runtimes like NVIDIA TensorRT LLM and serving frameworks like NVIDIA Dynamo, NVIDIA\u2019s <a href=\"https:\/\/blogs.nvidia.com\/blog\/inference-software-lowest-token-cost\/\" rel=\"nofollow noopener\" target=\"_blank\">software stack<\/a> is codesigned with the hardware to enable inference optimizations.<\/p>\n<p>While the results above reflect current Vera Rubin NVL72 performance, the full platform is a seven-chip architecture that also includes the NVIDIA Vera CPU, <a target=\"_blank\" href=\"https:\/\/developer.nvidia.com\/blog\/how-nvidia-groq-3-lpx-unlocks-ultrafast-interactivity-at-long-context-on-nvidia-vera-rubin\/\" rel=\"nofollow noopener\">Groq 3 LPU<\/a>, NVLink 6 Switch, BlueField-4 DPU, Spectrum-6 SPX and ConnectX-9 SuperNIC, all purpose-built for AI factories deploying agents at scale.<\/p>\n<p>Extreme codesign also extends to NVIDIA\u2019s co-engineering with its partners. Vera Rubin is in full production and is scaling across the ecosystem.\u00a0<\/p>\n<p>Learn more about the <a href=\"https:\/\/blogs.nvidia.com\/blog\/vera-rubin\/\" rel=\"nofollow noopener\" target=\"_blank\">NVIDIA Vera Rubin platform<\/a>.<\/p>\n","protected":false},"excerpt":{"rendered":"According to OpenRouter data, agentic AI workloads consume 15x more tokens than a simple chat request. Why?\u00a0 Consider&hellip;\n","protected":false},"author":2,"featured_media":149547,"comment_status":"","ping_status":"","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[2],"tags":[179,24,25,3887,26348,5243],"class_list":["post-149546","post","type-post","status-publish","format-standard","has-post-thumbnail","category-ai","tag-agentic-ai","tag-ai","tag-artificial-intelligence","tag-inference","tag-nvidia-vera-rubin","tag-think-smart"],"_links":{"self":[{"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/posts\/149546","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/users\/2"}],"replies":[{"embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/comments?post=149546"}],"version-history":[{"count":0,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/posts\/149546\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/media\/149547"}],"wp:attachment":[{"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/media?parent=149546"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/categories?post=149546"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/tags?post=149546"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}