NVIDIA Faces AI Inference Bottleneck—Vera Rubin and Groq 3 LPX Offer a Powerful Solution

NVIDIA’s Vera Rubin Gets a Major Inference Upgrade as Groq 3 LPX Hits Full Production

NVIDIA is making a major push into the rapidly growing world of agentic AI, announcing that its new Groq 3 LPX inference platform is now in full production and positioning it as a key extension of the company’s next-generation Vera Rubin AI infrastructure.

The move signals a bigger shift happening across the AI industry: the race is no longer just about training the biggest models. It is increasingly about how quickly, efficiently and cheaply those models can generate intelligence in real time.

And according to NVIDIA, the performance numbers are dramatic.

NVIDIA Claims 3,400 Tokens Per Second

In an Artificial Analysis benchmark using Gemma 4 31B, an open-source agentic AI model, NVIDIA said its Vera Rubin rack-scale system combined with Groq 3 LPX achieved 3,400 output tokens per second during 100,000-token long-context workloads.

That is particularly important for the next generation of AI agents.

Unlike traditional chatbots that simply answer a question, agentic AI systems can reason through problems, call tools, execute code, access data and communicate with other AI systems. These tasks require models to process enormous context windows while continuously generating new tokens.

NVIDIA says its system delivered performance four times faster than the nearest alternative platform in the benchmark.

The company is betting that this kind of token-generation speed will become one of the most important measurements in the AI infrastructure race.

AI Is Moving From Training to Reasoning

For years, the AI hardware race has been dominated by one question: How fast can companies train larger models?

That is beginning to change.

As reasoning models and autonomous AI agents become more capable, inference is becoming a massive computing challenge of its own. AI systems are generating longer responses, maintaining larger context windows and performing multiple steps before completing a task.

Even small delays can add up.

If an AI agent needs to generate thousands of tokens while calling multiple tools and interacting with other agents, slow token generation can quickly turn a responsive experience into a frustrating one.

NVIDIA’s answer is to redesign the AI factory around the entire workflow rather than relying on a single breakthrough chip.

Meet NVIDIA Groq 3 LPX

NVIDIA Groq 3 LPX is designed specifically for the latency-sensitive token generation required by agentic AI.

The architecture works alongside NVIDIA’s Vera Rubin NVL72 platform.

Rubin GPUs handle large-scale context processing, while LPX accelerators focus on the decode stage — the process of generating AI responses token by token.

The goal is to remove the traditional tradeoff between high throughput and low latency.

A rack-scale Groq 3 LPX deployment can reportedly include 256 LP30 accelerators, connected through direct chip-to-chip links and operating as a large-scale inference engine.

NVIDIA believes this combined GPU-and-LPU approach could be crucial as AI applications move toward real-time agents, coding assistants and multi-agent systems.

Nebius Becomes the First AI Cloud to Adopt Groq 3 LPX

NVIDIA also announced growing industry adoption of its Vera Rubin platform.

AI cloud company Nebius is becoming the first AI cloud provider to adopt NVIDIA Groq 3 LPX.

The technology will be integrated into the Nebius Token Factory, where NVIDIA says it will help developers build faster interactive AI agents, coding systems and other real-time AI applications.

The partnership is another sign that cloud infrastructure companies are preparing for an explosion in inference demand.

Training a frontier AI model may require enormous computing resources for a limited period. Serving that model to millions of users — and potentially billions of AI agents — could require a completely different kind of infrastructure.

CoreWeave Deploys Spectrum-X Multiplane

NVIDIA is also expanding the networking side of the AI factory.

CoreWeave has deployed NVIDIA Spectrum-X Multiplane in production, connecting Vera Rubin racks through multiple parallel network paths.

The technology is designed to create flatter and more resilient AI networks while avoiding the cost and complexity of adding another network tier.

According to NVIDIA, Spectrum-X Multiplane can scale to as many as 512,000 GPUs.

The company also claims that an eight-plane topology can maintain roughly 90% of total bandwidth even if one plane fails, with hardware recovery operating significantly faster than software-based approaches.

NVIDIA says the technology could deliver up to 1.6x higher AI factory output.

SpaceXAI Plans to Build Around NVIDIA Vera CPUs

Perhaps the most eye-catching announcement is NVIDIA’s expanding relationship with SpaceXAI.

The company plans to build and scale future AI architecture around NVIDIA Vera Rubin technology, including deployments ranging from terrestrial data centers to orbital satellites.

SpaceXAI plans to use NVIDIA Vera CPUs for CPU-heavy AI agent tasks, including:

  • Orchestration
  • Tool use
  • Code execution
  • Data processing
  • Simulation

The idea is simple: GPUs should spend as much time as possible performing the high-value AI compute they were designed for, while CPUs and specialized infrastructure handle the growing operational workload created by autonomous agents.

The AI Network Is Becoming Just as Important as the GPU

NVIDIA is also using the Vera Rubin rollout to emphasize a major reality of modern AI: compute is only part of the equation.

As AI clusters grow, moving data between GPUs becomes an enormous challenge.

NVIDIA Spectrum-X Ethernet is designed specifically for AI workloads, combining switches, SuperNICs and software into a unified networking platform.

The company says its Spectrum-X platform can provide 1.6x better AI networking performance compared with traditional off-the-shelf Ethernet.

Meanwhile, NVIDIA’s Spectrum-XGS technology is designed to connect multiple data centers, effectively allowing separate facilities to operate as one larger AI super-factory.

That could become increasingly important as AI companies struggle to find enough power, land and cooling capacity for gigantic new data centers.

NVIDIA Introduces Scale-In for Agentic AI Factories

NVIDIA also unveiled a new infrastructure category called Scale-In, powered by BlueField-4 processors and the NVIDIA DOCA software platform.

The goal is to accelerate the infrastructure services surrounding AI computing, including networking, storage, cybersecurity, provisioning and observability.

In other words, NVIDIA is no longer focusing only on making AI chips faster.

The company wants to accelerate the entire environment surrounding AI workloads.

As millions of users, applications and autonomous AI agents begin interacting with the same infrastructure, managing and securing those interactions could become just as important as raw GPU performance.

NVLink Fusion Opens NVIDIA’s AI Infrastructure to Custom Silicon

Another major part of NVIDIA’s strategy is NVLink Fusion.

The platform allows hyperscalers and AI-native companies to connect custom XPUs and CPUs to NVIDIA’s AI infrastructure.

That means companies can develop their own specialized processors while still using NVIDIA’s networking, rack architecture, cooling, power systems and software ecosystem.

NVLink Fusion includes NVIDIA’s sixth-generation NVLink technology, NVLink Switch and NVLink-C2C connectivity.

The strategy could give large cloud companies more flexibility while keeping NVIDIA at the center of the AI infrastructure ecosystem.

The Bigger Picture: NVIDIA Is Building a Token Factory

The most important takeaway from NVIDIA’s announcements may be the company’s evolving vision of the modern AI data center.

NVIDIA increasingly describes these massive computing systems as “token factories.”

The idea is that future AI infrastructure will be measured by how efficiently it converts massive amounts of computing power, data and energy into useful AI-generated tokens.

That means optimizing every layer simultaneously:

  • GPUs for large-scale AI computation
  • LPUs for ultrafast token generation
  • CPUs for orchestration and agent workloads
  • High-speed networking for massive AI clusters
  • Infrastructure processors for security and data services
  • Software for managing the entire system

NVIDIA calls this approach “extreme codesign.”

Instead of optimizing each component separately, the company is designing compute, networking, inference acceleration and infrastructure as one integrated system.

The AI Infrastructure Race Is Entering a New Phase

The Vera Rubin era appears to represent a major strategic shift for NVIDIA.

The company is preparing for a future where AI models don’t simply answer questions — they reason, plan, write code, operate software, collaborate with other agents and continuously process enormous amounts of information.

That future could require vastly more inference capacity than today’s AI systems.

With Groq 3 LPX now in full production, Spectrum-X Multiplane scaling massive GPU clusters, SpaceXAI planning to use Vera CPUs and Nebius becoming the first AI cloud to adopt the new inference platform, NVIDIA is making its next big bet clear.

The next AI war may not be won by the company with the fastest chip.

It may be won by the company that can build the fastest, most efficient and most scalable factory for generating intelligence.

And NVIDIA is determined to own every layer of that factory.

Subscribe

Explore More

Related Stories

Stay on op - Ge the daily news in your inbox