NVIDIA announced on August 24, 2026, that its Groq 3 LPX interactive AI inference accelerator has entered full production, marking the commercial arrival of technology from the company’s largest acquisition to date.

 

The Groq 3 LPX extends the NVIDIA Vera Rubin platform and is purpose-built to deliver ultrafast token generation for agentic AI systems. 

 

In independent benchmarking by Artificial Analysis, the accelerator achieved a record 3,400 output tokens per second running the open-source Gemma 4 31B model with a 100,000-token context window—the fastest performance recorded for that model.

 

Nebius, a leading AI cloud provider, will be the first to deploy the racks through its Nebius Token Factory inference platform, with systems expected to come online later this year. Groq itself is also listed among the earliest planned adopters.

 

Agentic workloads generate massive volumes of tokens across hundreds or thousands of sequential inference steps as models reason, call tools, inspect files, write and test code, and iterate. 

 

Faster generation rates directly determine how responsive these systems feel to users. NVIDIA states that Groq 3 LPX provides roughly 4x faster responsiveness for agents and other latency-sensitive workloads compared with the nearest alternative platform.

 

“Inference is the growth engine of AI,” said Jensen Huang, founder and CEO of NVIDIA. “Vera Rubin extends that vision with workload-optimized AI factory configurations designed for the era of agentic AI, advancing the performance frontier with LPX for ultrafast token generation. 

 

This transforms how intelligence is produced, delivering another giant leap in AI throughput, efficiency and responsiveness, just as demand for AI computation is accelerating worldwide.”

 

The architecture addresses two distinct challenges of agentic AI: efficiently processing large amounts of context and generating tokens with extremely low latency. 

 

Vera Rubin NVL72 systems handle the broader training and inference platform, while Groq 3 LPX accelerates the generation (decode) phase that determines interactivity for individual users.

 

NVIDIA packages 256 individual Groq 3 chips into each LPX rack. The underlying Groq design includes 500 megabytes of on-die SRAM to reduce memory bottlenecks.

 

The chips are manufactured by Samsung, while NVIDIA’s GPUs continue to be produced by TSMC. The LPX racks are designed to work alongside Vera CPUs, Rubin GPUs, BlueField-4 DPUs, Vera BlueField-4 STX storage and Spectrum-6 SPX Ethernet in rack-scale configurations.

 

Danila Shtan, chief technology officer of Nebius, said: “Generation is the phase of inference that determines how responsive an AI system actually is, and that’s exactly what NVIDIA Groq 3 LPX is built to accelerate. 

 

As the first AI cloud bringing it to production via Nebius Token Factory, we’re making sure every step of an agent’s loop feels instant—through the same API developers are already using, with no migration to a new stack.”

 

The move commercializes technology acquired when NVIDIA bought assets from Groq in a deal valued at approximately $20 billion, the largest in the company’s history.

 

At the March 2026 GTC unveiling of Vera Rubin and Groq 3 LPX, Huang projected cumulative sales of $1 trillion across Blackwell and Vera Rubin systems through 2027 and indicated that a portion of data-center capacity intended for coding applications would be allocated to the specialized inference accelerators.

 

Low-latency inference chips such as the Groq 3 LPX do not replace GPUs. They complement them by optimizing the decode phase of serving, allowing cloud providers to offer premium, latency-sensitive service tiers. 

 

Competitors including AMD have pursued similar specialized inference partnerships, and OpenAI has previewed ultrafast modes powered by alternative architectures.

 

For AI factories and neoclouds serving high-volume, real-time agent workloads—coding assistants, multi-step reasoning systems, tool-using agents and interactive applications—the availability of production Groq 3 LPX racks removes a key bottleneck.

 

Developers gain access to extreme token generation speeds without changing existing APIs or stacks when using Nebius Token Factory.

 

NVIDIA is currently ramping shipments of Vera Rubin systems that entered production earlier in 2026. The addition of dedicated high-speed generation hardware completes a codesigned stack spanning seven chips and multiple purpose-built rack configurations optimized for the dual demands of context processing and rapid sequential generation.

 

As agentic AI moves from research demonstrations into production use cases, infrastructure that can keep pace with long-horizon, multi-step workflows becomes a competitive differentiator. 

 

The full-production status of Groq 3 LPX, combined with first-cloud deployment by Nebius later this year, positions the technology as a practical building block for the next wave of responsive, tool-using AI systems.

 

NVIDIA reports earnings later this week, providing the market with the next data point on how quickly demand for these specialized inference platforms is translating into revenue and deployment volume.