Nvidia has put its Groq 3 LPX interactive inference accelerator into full production, the company announced on August 24, 2026, during the Hot Chips conference. The chip, an extension of Nvidia's Vera Rubin platform, delivered a record 3,400 output tokens per second in Artificial Analysis benchmarking while running the open-source Gemma 4 31B model with an active 100,000-token context window — the fastest performance ever recorded for that model.
The launch marks the most significant product milestone yet from Nvidia's $20 billion acquisition of inference specialist Groq, a deal the company is now moving to capitalize on as demand for low-latency inference surges. CNBC reported that Nvidia says the first Groq racks will be online before the end of the year. For more context on this story, see our ongoing AI industry coverage.
Why Token Generation Speed Matters
Agentic AI systems — programs that reason, call tools, write and test code, and iterate across hundreds or thousands of inference steps — generate enormous volumes of output tokens. The speed at which those tokens are produced, known as interactivity or decode throughput, determines how quickly an agent can complete each step of its work.
Nvidia says Groq 3 LPX provides 4x faster responsiveness for agents and latency-sensitive workloads than the nearest alternative platform. In heterogeneous deployments, Vera Rubin NVL72 systems handle context ingestion and prefill computation, while Groq 3 LPX units process the latency-sensitive token generation phase.
"Inference is the growth engine of AI," Jensen Huang, founder and CEO of Nvidia, said in the announcement. "Vera Rubin extends that vision with workload-optimized AI factory configurations designed for the era of agentic AI, advancing the performance frontier with LPX for ultrafast token generation."
Inside the Hardware
According to StorageReview's technical analysis, each 2U liquid-cooled compute tray carries sixteen Groq 3 LPUs alongside a host CPU, fabric expansion logic, and DRAM, plus either a BlueField-4 DPU or a ConnectX-9 network card. The tray features 32 LPU chip-to-chip optical links and 400Gb/s Ethernet on the front panel.
At rack scale, a Groq 3 LPX deployment incorporates up to 256 LP30 accelerators linked through high-speed, direct chip-to-chip interconnects. The racks slot into Nvidia's broader Vera Rubin AI factory infrastructure, which the company describes as its most extensive platform ever: seven discrete chips across five purpose-built rack configurations, including BlueField-4 DPUs, Vera CPU racks, Vera BlueField-4 STX storage, and Spectrum-6 SPX Ethernet networking.
Nvidia also refreshed its platform-level performance claims, stating that Vera Rubin NVL72 delivers up to 30 times the token throughput per megawatt compared with GB300 NVL72 on a 140,000-token agentic coding workload, with the advantage widening as per-user interactivity targets rise.
A Benchmark With an Asterisk
The headline 3,400 tokens-per-second figure comes with a caveat worth noting. As StorageReview pointed out, Nvidia's result was measured on a private pre-release endpoint on August 21, while the competing numbers it was compared against come from live serverless production endpoints. Real-world performance on generally available services may differ from the record-setting benchmark configuration.
Still, independent benchmarking organization Artificial Analysis recorded the result as the fastest ever measured for Gemma 4 31B at that context length, and the underlying claim — that purpose-built silicon can dramatically accelerate the decode phase of inference — aligns with how the industry has been re-architecting for agentic workloads.
Cloud Adoption Begins
Nebius, the Amsterdam-headquartered AI cloud, is the first provider to commit to deploying Groq 3 LPX, planning to bring it into Nebius Token Factory, its production inference platform.
"Generation is the phase of inference that determines how responsive an AI system actually is, and that's exactly what NVIDIA Groq 3 LPX is built to accelerate," said Danila Shtan, chief technology officer of Nebius. "As the first AI cloud bringing it to production via Nebius Token Factory, we're making sure every step of an agent's loop feels instant — through the same API developers are already using, with no migration to a new stack."
Inference provider Groq, now part of the Nvidia family, plans to be among the platform's earliest adopters.
The Bigger Picture
The full-production announcement signals how sharply the AI infrastructure market has pivoted from training to inference. As enterprises shift budgets from raw model pre-training toward real-time reasoning and autonomous agent orchestration, the economics of a token — how fast it can be generated and at what energy cost — have become the central competitive battleground.
For Nvidia, Groq 3 LPX extends a defense of its AI factory franchise at a moment when custom silicon from cloud providers and startups alike is chasing the same agentic inference workloads. Racks coming online this year, as CNBC reported, would put the first production systems in customers' hands months after the acquisition closed — a fast integration by semiconductor industry standards.
The company did not disclose pricing for the new systems. Groq 3 LPX will be offered on a when-and-if-available basis through AI cloud partners, with Nebius expected to be the first to offer access to developers through its existing API endpoints.
---
Stay Ahead of AIGet the latest AI news, analysis, and breakthroughs — all in one place.
Read more AI news →