NVIDIA Says Vera Rubin Raises Agentic AI Throughput per Megawatt
NVIDIA published AgentX benchmark results claiming Vera Rubin NVL72 delivers major throughput-per-megawatt gains for production-style agentic AI workloads compared with prior generations.
What NVIDIA reported
On August 24, 2026, NVIDIA published new performance-per-watt results for agentic AI workloads on Vera Rubin and Blackwell systems. The company used AgentX, an open-source benchmark in the InferenceX suite that replays production-style coding-agent sessions with long-context prefill, KV-cache reuse, tool-call gaps and changing concurrency.
The headline efficiency claim
NVIDIA says Vera Rubin NVL72 achieved up to 30 times higher AI-factory throughput per megawatt than GB300 NVL72 in the AgentX workloads it reported. It also says GB300 NVL72 extends earlier generational efficiency gains over H200 NVL8, with especially large improvements on very large mixture-of-experts models. These are vendor-published benchmark results, so workload selection, software tuning and independent reproduction remain important context.
Why agentic inference is different
Agent workloads do not behave like a simple stream of independent prompts. They may repeatedly reuse context, pause for tools, spawn subagents, switch between prefill-heavy and decode-heavy phases, and operate under variable concurrency. Benchmarks that reproduce those behaviors can provide a more realistic view of infrastructure efficiency than single-request token-throughput tests alone.
The system-level stack
NVIDIA attributes the reported gains to more than accelerator silicon. The stack includes mixture-of-experts serving runtimes such as SGLang, TensorRT-LLM and vLLM, optimized kernels, lower-precision formats including MXFP4 and MXFP8, NVIDIA Dynamo for session-aware serving, and NVLink scale-up fabric across rack-scale systems. The result reinforces the industry's shift toward co-optimizing models, runtimes, networking and hardware together.
What to watch next
Independent AgentX runs, broader model coverage and real production power measurements will be the most useful follow-up evidence. If the efficiency gains generalize, power-constrained data centers could support substantially more long-running agent work within the same electrical envelope, making performance per megawatt a central competitive metric for AI infrastructure.
This article is built from the source material below. Open the originals for full context and the latest updates.