Editorial illustration for NVIDIA's Vera Rubin NVL72 Leads MLPerf Inference Debut
NVIDIA Vera Rubin NVL72 Dominates MLPerf Inference
NVIDIA's next-generation Vera Rubin NVL72 system posted its first MLPerf Inference numbers on Wednesday, and the results came with a specific figure attached: up to 3.7x better throughput than the current GB300 NVL72 rack. That's the headline from MLPerf Inference v6.1, the latest round of results from the industry's standard AI benchmark suite, but it's not the only data point NVIDIA is pointing to.
The company also submitted a 288-GPU run spanning four GB300 NVL72 racks that hit 99% scaling efficiency, meaning throughput grew almost in lockstep as GPUs were added rather than tapering off. Separately, software updates baked into the v6.1 submission pushed performance up to 1.6x higher than what the same hardware managed in v6.0, with more gains logged after the submission window closed.
Those three threads, raw throughput on a new system, near-linear scaling on an existing one, and software gains squeezed from unchanged silicon, map onto how buyers actually judge inference hardware: what it can generate, how cheaply that output scales as racks multiply, and how much more value shows up over time without new purchases. NVIDIA ran its debut Vera Rubin numbers on two of the suite's toughest workloads, DeepSeek-R1 and Qwen3-VL.
Vera Rubin NVL72 delivers up to 3.7x higher throughput than GB300 NVL72 on Qwen3-VL across offline, server and interactive scenarios, using vLLM with the NVIDIA Dynamo open source inference framework. On DeepSeek-R1, using the NVIDIA TensorRT-LLM library, throughput is up to 2.5x higher than GB300 NVL72.
Why this matters
NVIDIA's numbers on DeepSeek-R1 and Qwen3-VL are impressive on paper, but "preview results" is the key phrase here. Vera Rubin NVL72 hasn't shipped, and MLPerf's own rules let vendors submit hardware that customers can't yet buy. A 3.7x jump over GB300 NVL72 tells us NVIDIA's roadmap is intact and its engineers know how to game a benchmark suite that rewards throughput above almost everything else.
For developers budgeting inference costs and founders pricing token-based products, the real question isn't peak throughput, it's what happens when Vera Rubin ships in volume, how software optimizations translate outside NVIDIA's own test harness, and whether "platform fungibility" holds up when you're not running NVIDIA's reference stack. Researchers comparing architectures should treat this as a preview of NVIDIA's marketing, not a settled comparison. Worth watching: independent benchmarks once actual silicon reaches customers, and how AMD or custom silicon vendors respond on the same DeepSeek-R1 and Qwen3-VL workloads once they're not just chasing a headline number NVIDIA set first.
Common Questions Answered
What performance improvements does the Vera Rubin NVL72 deliver compared to GB300 NVL72 on Qwen3-VL?
The Vera Rubin NVL72 delivers up to 3.7x higher throughput than GB300 NVL72 on Qwen3-VL across offline, server and interactive scenarios. This performance gain is achieved using vLLM with the NVIDIA Dynamo open source inference framework, representing a significant advancement in NVIDIA's inference capabilities.
How does Vera Rubin NVL72 perform on DeepSeek-R1 according to MLPerf Inference v6.1 results?
On DeepSeek-R1, the Vera Rubin NVL72 achieves up to 2.5x higher throughput than GB300 NVL72 using the NVIDIA TensorRT-LLM library. These results demonstrate strong performance improvements across different model architectures and inference frameworks.
What scaling efficiency did NVIDIA achieve with a 288-GPU Vera Rubin NVL72 configuration?
NVIDIA submitted a 288-GPU run spanning four GB300 NVL72 racks that achieved 99% scaling efficiency in the MLPerf Inference v6.1 benchmark. This near-perfect scaling indicates excellent performance optimization across multiple racks in the system.
Why are Vera Rubin NVL72's MLPerf results described as 'preview results'?
The Vera Rubin NVL72 hasn't shipped yet, and MLPerf's rules allow vendors to submit hardware that customers cannot yet purchase. This means the benchmark results represent NVIDIA's roadmap and engineering capabilities rather than performance from systems currently available in the market.
What is the significance of Vera Rubin NVL72's performance for developers and AI product founders?
The performance improvements demonstrated by Vera Rubin NVL72 are important for developers budgeting inference costs and founders pricing token-based products, as they indicate the trajectory of NVIDIA's inference capabilities and potential cost efficiencies. The results show that NVIDIA's engineers have optimized the system to excel at throughput, which directly impacts the economics of AI inference workloads.
Further Reading
- NVIDIA Vera Rubin NVL72 Delivers Leading Performance in MLPerf Inference v6.1 Debut - NVIDIA Blog
- AI Infra Summit: NVIDIA Vera Rubin and DSX Platform Deliver Higher Tokens per Watt - NVIDIA Blog
- NVIDIA Vera Rubin and Blackwell Set a New Standard for Agentic AI Performance Per Watt - NVIDIA Developer Blog
- NVIDIA Vera Rubin NVL72 Rack at Hot Chips 2026 - ServeTheHome
- Rubin NVL72 Agentic Inference: 67x better Performance ... - SemiAnalysis