NVIDIA Vera Rubin Delivers 30x More Agentic AI Throughput Per MW

NVIDIA Claims Vera Rubin Could Deliver 30x More AI Agent Work Per Megawatt

The AI race is no longer just about building smarter models. It is increasingly about how efficiently the infrastructure behind those models can handle an explosion of tokens.

And NVIDIA believes its upcoming Vera Rubin NVL72 platform could deliver a massive leap.

According to new performance data released by NVIDIA, Vera Rubin NVL72 systems can deliver up to 30x higher throughput per megawatt on agentic AI workloads compared with GB300 NVL72 systems.

The company also claims the new platform could reduce the cost per million tokens by as much as 35x.

If those numbers hold up in broader testing, the implications for the rapidly growing AI industry could be enormous.

AI Agents Are Creating a Token Explosion

A simple chatbot question might involve only a relatively small amount of context.

AI agents are a completely different story.

According to data cited from OpenRouter, agentic AI workloads can consume 15 times more tokens than a standard chat request.

The reason is simple: AI agents do not just generate an answer and stop.

Imagine an AI system researching a company before making an investment recommendation.

The agent may search financial databases, scan news reports and regulatory filings, launch additional sub-agents to compare competitors, run valuation models and then combine all of that information into a final recommendation.

Every step generates more context.

That growing context becomes part of the next request, creating enormous computational demands as the AI continues reasoning, using tools and spawning additional agents.

The same pattern is emerging in software development, customer support and deep research.

As AI agents move from demonstrations into real-world production, the infrastructure powering them is facing an entirely new challenge.

How do you process an exploding number of tokens without sending power consumption and operating costs through the roof?

NVIDIA Says Vera Rubin Changes the Equation

NVIDIA’s answer is the Vera Rubin NVL72.

Using the SemiAnalysis AgentX workload, which includes recorded real-world agentic coding sessions with context growth, tool calls and sub-agent activity preserved, NVIDIA measured the performance of its next-generation platform against GB300 NVL72.

The result?

Up to 30x higher throughput per megawatt.

In practical terms, NVIDIA says an AI factory constrained by the same power budget could potentially run dramatically more agentic workloads.

The company also says Vera Rubin NVL72 could deliver up to 35x lower cost per million tokens compared with GB300 NVL72.

That matters because the economics of AI are increasingly tied to two critical numbers:

  • How much AI work can be produced from every megawatt of power.
  • How much it costs to generate every million tokens.

For companies operating enormous AI factories, even a major improvement in either number could have a huge impact on revenue and profitability.

Why AI Agents Are So Much Harder to Run

Traditional AI benchmarks often focus on relatively straightforward requests.

A user enters a prompt.

The model processes it.

The model generates an answer.

Agentic AI does not work that way.

An agent may operate across dozens or even hundreds of steps. Context can accumulate into hundreds of thousands of tokens, while input and output sizes constantly change throughout the task.

That means measuring a single inference request may not accurately represent real-world agent performance.

NVIDIA’s new benchmark data attempts to capture the entire workflow, including the long context, tool calls and sub-agent activity that increasingly define modern AI systems.

The SemiAnalysis AgentX workload includes real-world coding trajectories and is designed to reflect how AI agents actually behave when working through complex tasks.

NVIDIA says the Blackwell platform has already demonstrated strong performance across models including Kimi K3, MiniMax M3, GLM 5.3, Qwen 3.5 and DeepSeek V4 Pro.

GB300 NVL72 previously showed up to 15x better throughput per megawatt than NVIDIA Hopper on the DeepSeek V4 Pro model, according to NVIDIA.

Now, Vera Rubin is being positioned as the next major leap.

The Secret Is Extreme AI Codesign

NVIDIA says the performance gains are not coming from a single chip upgrade.

Instead, Vera Rubin NVL72 is designed around what the company calls extreme codesign.

The platform combines hardware, networking and software optimizations specifically aimed at the complex demands of AI agents.

Key technologies include:

  • Disaggregated serving, which separates context processing from response generation.
  • Rate matching, designed to keep different parts of the inference system operating efficiently.
  • Large-scale expert parallelism for mixture-of-experts AI models.
  • Distributed KV caching, allowing previously processed context to remain accessible across the GPU system.
  • KV-aware routing, reducing unnecessary recomputation during long AI sessions.
  • Fused CUDA kernels, which combine operations to keep GPUs busy rather than waiting for data.

The Rubin GPU architecture also introduces enhanced fifth-generation Tensor Cores and a third-generation Transformer Engine.

Meanwhile, NVFP4 quantization reduces the memory required for model weights while aiming to increase throughput without significantly affecting output quality.

A Seven-Chip AI Factory Built for Agents

Vera Rubin is more than just a GPU platform.

The full architecture includes multiple components designed to work together inside future AI factories.

That includes the NVIDIA Vera CPU, Groq 3 LPU, NVLink 6 Switch, BlueField-4 DPU, Spectrum-6 SPX and ConnectX-9 SuperNIC.

NVIDIA’s strategy is clear.

Instead of treating AI hardware, networking and software as separate pieces, the company is designing the entire system as one massive computing architecture optimized for increasingly complex AI workloads.

The NVL72 scale-up domain allows GPUs to communicate using NVIDIA’s NVLink technology, providing the high-bandwidth and low-latency connections required for techniques such as expert parallelism and distributed memory systems.

NVIDIA says its latest NVLink and NVLink Switch technology delivers dramatically higher packet rates and lower latency compared with conventional networking alternatives.

The AI Infrastructure War Is Entering a New Phase

For years, the AI hardware race has largely focused on who could build the fastest accelerator.

That is changing.

The next battle may be about tokens per watt.

AI companies are racing to deploy increasingly powerful agents capable of coding, researching, analyzing data and performing multi-step tasks autonomously.

But those capabilities come with a price.

More reasoning.

More context.

More tool calls.

More sub-agents.

And, ultimately, far more tokens.

That makes power efficiency one of the biggest challenges facing the future of artificial intelligence.

NVIDIA’s early Vera Rubin results suggest the company is preparing for that future with an architecture built specifically around agentic AI at enormous scale.

The company’s latest figures are based on early measurements and are currently pending SemiAnalysis review, while continued software optimization could further change performance over time.

Still, the headline number is impossible to ignore.

If Vera Rubin can truly deliver up to 30x more agentic AI throughput from the same megawatt of power, the economics of running AI agents at massive scale could be about to change dramatically.

The AI race may no longer be won simply by whoever builds the smartest model.

It could be won by whoever can produce the most intelligence for every watt of power.

Subscribe

Explore More

Related Stories

Stay on op - Ge the daily news in your inbox