Skip to main content
NVIDIA Blackwell GPU architecture, a powerful AI agent, excels in SemiAnalysis benchmark performance.

Editorial illustration for NVIDIA Blackwell Leads AI Agent Performance in SemiAnalysis Benchmark

NVIDIA Blackwell Dominates AI Agent Benchmark

4 min read

An AI agent asked to evaluate a company for investment doesn't just answer a question. It queries financial databases, pulls news and filings, spins up a sub-agent to run peer comparisons, models valuations, then stitches all of it into a recommendation. Every one of those steps feeds tokens into the next, and OpenRouter data shows that pattern pushes agentic workloads to consume 15 times more tokens than a simple chat request. Software development, customer service, deep research: the same loop of reasoning, tool calls, and sub-agent spawning shows up everywhere agentic AI gets deployed.

That token appetite is now a hardware problem. As agentic systems move from demos into production, the infrastructure underneath them has to keep pace with demand that looks nothing like a chatbot query. NVIDIA says its Vera Rubin NVL72 platform was built with that shift in mind, and new measured results, using SemiAnalysis's AgentX benchmark built from recorded real-world coding sessions, put a number on the gap between generations. The comparison centers on NVIDIA's own GB300 NVL72 systems, and the gains show up specifically in throughput per megawatt, the metric that matters most for power-constrained data centers.

New measured performance data shows NVIDIA Vera Rubin NVL72 systems deliver up to 30x higher throughput per megawatt than NVIDIA GB300 NVL72 on agentic workloads.

Why this matters

For anyone building or budgeting agentic systems, the token math here is the real story. If agents already burn 15x the tokens of a basic chat request, and that multiplies further with sub-agents chaining research, comparison, and synthesis steps, throughput-per-watt stops being an infrastructure footnote and becomes a line item founders have to plan around. SemiAnalysis putting GB300 NVL72 at 15x Hopper's efficiency, with Vera Rubin NVL72 claiming up to 30x, is NVIDIA's benchmark to promote, and we'd want independent runs before treating those numbers as settled.

But the direction is worth watching closely: if agentic workloads are the growth driver vendors say they are, efficiency-per-watt becomes the metric that decides which stack you build on, not raw FLOPS. Teams running Kimi K3, MiniMax M3, or DeepSeek V4 Pro in production should be asking their cloud providers which hardware generation is actually serving those models, because the gap between architectures, if it holds up under scrutiny, will show up directly in your compute bill.

Common Questions Answered

Why do AI agents consume 15 times more tokens than simple chat requests?

AI agents consume significantly more tokens because they execute multi-step workflows that feed tokens into each successive step. For example, an investment evaluation agent queries financial databases, pulls news and filings, spins up sub-agents for peer comparisons, models valuations, and stitches results together, with each step generating and consuming additional tokens in the process.

How much more efficient is NVIDIA Vera Rubin NVL72 compared to GB300 NVL72 for agentic workloads?

According to measured performance data, NVIDIA Vera Rubin NVL72 systems deliver up to 30x higher throughput per megawatt than NVIDIA GB300 NVL72 on agentic workloads. This represents a substantial efficiency improvement for handling the token-intensive demands of AI agent applications.

What types of applications beyond investment analysis benefit from agentic AI systems?

Agentic AI systems with their multi-step token-consuming workflows are beneficial for software development, customer service, and deep research applications. These domains all follow the same looping pattern where agents query multiple data sources, perform analysis, and synthesize recommendations.

Why is throughput-per-watt efficiency critical for founders building agentic systems?

Since agents already consume 15 times more tokens than basic chat requests, and this multiplies further when sub-agents chain research, comparison, and synthesis steps together, throughput-per-watt efficiency directly impacts operational costs and becomes a significant line item that founders must budget and plan around. Infrastructure efficiency is no longer just a technical footnote but a core business consideration.

LIVE23:05Google's ME-POIs Adds "How a Place Is Used" to POI Embeddings