MLCommons released the MLPerf Inference v6.1 benchmark results on September 16, 2026, and the headline belonged to NVIDIA's next-generation platform. In its first-ever preview submission, the Vera Rubin NVL72 system delivered up to 3.7 times the throughput of NVIDIA's current GB300 NVL72 flagship, according to NVIDIA's announcement, an early signal of what the successor to Blackwell Ultra will mean for AI inference economics.
Vera Rubin is NVIDIA's next data-center platform, named for the astrophysicist who provided key evidence for dark matter. The MLPerf preview results are the first public, independently organized performance data for the system. For more context on this story, see our ongoing breaking AI news.
The Numbers Behind the Debut
NVIDIA submitted Vera Rubin NVL72 preview results on two of the most demanding workloads in the v6.1 suite: the DeepSeek-R1 reasoning model and the Qwen3-VL vision-language model.
On Qwen3-VL, Vera Rubin NVL72 delivered up to 3.7 times higher throughput than GB300 NVL72 across offline, server and interactive scenarios, running vLLM with NVIDIA's Dynamo open-source inference framework. On DeepSeek-R1, using NVIDIA's TensorRT-LLM library, throughput was up to 2.5 times higher than the previous flagship.
Both figures come from NVIDIA's own submissions to MLCommons' Closed Division, meaning they were produced under MLPerf's audited rules, though preview results are explicitly early measurements of a platform still being optimized.
How the Gains Are Built
NVIDIA attributes the jump to full-stack co-design across hardware and software rather than any single component.
On the silicon side, Vera Rubin brings enhanced Tensor Cores and an updated Transformer Engine that accelerate both stages of inference — the compute-heavy prefill phase and the memory-bound decode phase. The platform leans on NVFP4 precision, which shrinks the memory footprint of model weights, attention and the KV cache, raising throughput with what NVIDIA describes as minimal loss of output quality.
On the software side, the submissions used disaggregated serving, which separates prefill and decode onto different tasks, along with large-scale expert parallelism to keep the mixture-of-experts layers in models like DeepSeek-R1 and Qwen3-VL fully fed with work.
The interconnect is the third pillar. The NVL72 rack design ties 72 GPUs into a single scale-up domain using sixth-generation NVLink and NVLink Switch, which NVIDIA says delivers 10 times higher packet rates and 3 times lower latency than off-the-shelf Ethernet — the foundation that makes rack-scale disaggregated serving practical.
GB300 Shows Near-Perfect Scaling
The current flagship had its own news in the v6.1 round. NVIDIA's DeepSeek-R1 submission scaled from a single GB300 NVL72 rack — 72 GPUs — to four racks, 288 GPUs, achieving 99 percent scaling efficiency in the offline scenario, with throughput growing nearly in proportion to the hardware added.
Scaling efficiency is a metric data-center operators watch closely because it determines whether buying more hardware actually buys proportionally more capacity. Near-linear scaling across racks suggests the bottleneck at this scale remains per-GPU performance rather than inter-rack communication.
Software Keeps Compounding
A recurring pattern in MLPerf rounds held again: software improvements alone delivered up to 1.6 times higher performance in NVIDIA's v6.1 submissions compared with v6.0, and the company says optimizations continued to land after the v6.1 submission deadline, delivering further gains.
For buyers, the implication is that a given rack gets faster over its lifetime without a hardware refresh — a dynamic NVIDIA has used to argue that its installed base appreciates rather than depreciates as the software stack matures.
Agentic Workloads Reshape the Benchmarks
The v6.1 round also reflected a shift in what inference is being used for. NVIDIA pointed to preview testing on SemiAnalysis's AgentX benchmark, designed around AI agents that reason, plan and act across multiple steps, where it says Vera Rubin NVL72 delivered 30 times the performance of GB300 NVL72.
MLCommons is also preparing a dedicated benchmark for this shift: the upcoming MLPerf Endpoints benchmark aims to standardize measurement of agentic inference workloads that traditional throughput tests capture poorly. As inference moves from single-prompt chat toward long multi-step agent sessions, latency and interactive consistency matter as much as raw tokens per second — and benchmarks are being rebuilt accordingly.
Nebius and the Ecosystem Arrive Early
Cloud provider Nebius also submitted Vera Rubin NVL72 preview results, which NVIDIA highlighted as evidence of strong early ecosystem performance. Early third-party submissions are typically a sign that next-generation systems are moving from lab sampling toward broader availability, though neither NVIDIA's post nor MLCommons' release specifies general-availability timing.
Why It Matters for AI Economics
NVIDIA frames the results in revenue terms: each Vera Rubin NVL72 rack generates significantly more tokens, serves more users and lowers cost per token relative to a GB300 NVL72 rack. In an industry where inference demand is being driven by reasoning models and agent workloads that consume orders of magnitude more compute per query than simple chat, per-rack throughput is the number that determines how many users a fixed data-center budget can serve.
The usual caveats apply: these are vendor submissions in a preview round, on two models chosen by NVIDIA, with optimizations still landing. But the direction is unambiguous. MLPerf's first look at Vera Rubin suggests the performance climb that has defined AI infrastructure since 2023 has no plateau in sight — and that the next rack-refresh cycle will again reset the economics of serving AI at scale.
---
Stay Ahead of AIGet the latest AI news, analysis, and breakthroughs — all in one place.
Read more AI news →