Recent findings from OpenRouter indicate that agentic AI workloads demand a staggering 15 times more tokens compared to standard chat requests. What accounts for this significant difference?
Imagine an AI agent tasked with investigating a company for a potential investment. The agent must access financial databases, analyze relevant news and filings, utilize a sub-agent for peer comparisons, and model valuations—all while synthesizing this information into a cohesive recommendation. During this process, these agents and their sub-agents engage in continuous reasoning until the objective is fulfilled, creating a growing need for tokens. As each step builds on the last, effective handling of extensive context becomes crucial for the performance of agentic AI.
This escalation in token consumption is consistent across various applications of agentic AI, whether in software development, customer support, or in-depth research.
As the adoption of agentic AI expands across different sectors, it is vital for the underlying infrastructure to efficiently cater to this rising token demand.
New performance metrics reveal that NVIDIA’s Vera Rubin NVL72 systems provide up to 30 times greater throughput per megawatt compared to the NVIDIA GB300 NVL72 when handling agentic workloads. This inference throughput was assessed using the SemiAnalysis AgentX workload, which involved recorded real-world sessions of agentic coding, maintaining actual context growth, tool interactions, and sub-agent activations. For energy-limited AI factories, this improvement translates directly to a 30-fold increase in agentic tasks executed for the same energy consumption.
These preliminary findings regarding the Vera Rubin NVL72 highlight NVIDIA’s rapid advancements in innovation. Ongoing software enhancements promise to boost performance further across both the Vera Rubin NVL72 and GB300 NVL72 systems.
When discussing agentic workloads, the nature of the work is distinctly different from that of chat or document summarization, where the input-output sequences typically range from 1,000 to 8,000 tokens. In contrast, agentic sessions can see context accumulate through multiple steps, often resulting in hundreds of thousands of input tokens, with variations in both input and output lengths. Thus, performance measurement strategies need to evolve to reflect the entirety of the agent workflow, rather than focusing on a single inference request.
The following performance results are reflective of real-world coding scenarios with agentic AI.
The NVIDIA Blackwell platform showcases superior performance across various agentic models, including Kimi K3, MiniMax M3, GLM5.3, Qwen3.5, and DeepSeek V4 Pro. For instance, the GB300 NVL72 exhibits up to 15 times superior throughput per megawatt compared to the NVIDIA Hopper architecture when utilizing the DeepSeek V4 Pro model. This enhancement stems from the benefits of a larger GPU domain and code-signed software that significantly boosts inference efficiency.
Vera Rubin enhances this advantage, providing as much as 30 times higher throughput per megawatt compared to GB300 NVL72 on the same DeepSeek V4 Pro model. These initial results, derived from the SemiAnalysis AgentX workload and currently pending review from SemiAnalysis, do not yet account for Vera CPU performance in tool interactions.
Through technologies such as NVIDIA's DSX MaxLPS, power management across GPU, rack, and workload levels enables the provisioning of up to 40% more GPUs under the same power budget, pushing throughput per megawatt to new heights in large-scale AI environments.
Furthermore, the relationship between throughput per megawatt and the cost of token production is significant. The Vera Rubin NVL72 achieves costs that are up to 35 times lower per million tokens compared to GB300 NVL72, allowing agents to operate continuously at scale throughout the entirety of their customers' workloads.
In power-limited AI factories, the throughput rate directly influences revenue, while the cost per million tokens dictates profit margins.
To achieve optimal performance for agentic AI, advanced inference optimizations are crucial. NVIDIA’s Vera Rubin NVL72 facilitates a wide array of such optimizations, thanks to a highly integrated design across all layers of the platform.
Disaggregated serving separates context processing from response generation, allowing each function to scale independently, while rate matching coordinates the token production speed between prefill and decode GPUs to maximize overall efficiency. Large-scale expert parallelism distributes expert sub-networks across the GPU domain, and distributed KV-caching expands memory access, ensuring previously processed context is readily available without needing to be recomputed.
KV-aware routing directs incoming requests to GPUs that have relevant cached context, minimizing redundant calculations during extensive sessions. Additionally, fused CUDA kernels like MegaMoE streamline multiple computation and communication tasks into single processing passes, ensuring GPUs remain active rather than idle.
The enhanced fifth-generation Tensor Cores and third-generation Transformer Engine in NVIDIA Rubin GPUs accelerate the inference process significantly. Through NVFP4 quantization, model weights are compressed to 4-bit precision, optimizing both memory use and throughput while maintaining output quality.
The NVL72 scale-up domain, integral to the architecture of both Vera Rubin and Grace Blackwell, facilitates high-bandwidth, low-latency communication between GPUs—essential for methods like large-scale expert parallelism and distributed KV-caching. The purpose-built NVIDIA NVLink technology and sixth-generation NVLink Switches provide tenfold higher packet rates with threefold lower latency compared to standard Ethernet options.
NVIDIA's software ecosystem—including optimized CUDA kernels, TensorRT LLM inference runtimes, and NVIDIA Dynamo serving frameworks—is meticulously designed alongside the hardware to enhance inference operations.
While the current results illustrate the capabilities of the Vera Rubin NVL72, the comprehensive platform encompasses a seven-chip architecture that includes elements such as the NVIDIA Vera CPU, Groq 3 LPU, NVLink 6 Switch, BlueField-4 DPU, Spectrum-6 SPX, and ConnectX-9 SuperNIC, all specifically engineered for deploying agentic workflows at scale.
NVIDIA continues to prioritize extreme codesign principles in collaboration with its partners, ensuring that the Vera Rubin system is fully operational and scaling within the tech ecosystem.




