AiPhreaks ← Back to News Feed

NVIDIA Vera Rubin and Blackwell Set a New Standard for Agentic AI Performance per Watt

By Jakub Antkiewicz

2026-08-25T08:41:10Z

NVIDIA Vera Rubin Achieves 30x Efficiency Gain in New Agentic AI Benchmark

NVIDIA has released preview results for its upcoming Vera Rubin platform, showing a substantial increase in efficiency for agentic AI workloads. According to data from the SemiAnalysis AgentX benchmark, the Vera Rubin NVL72 system delivers up to 30 times higher AI throughput per megawatt than the current-generation Blackwell GB300 NVL72. These findings are significant as the industry shifts from simple, single-turn inference to complex, multi-step agentic workflows, which now represent a major driver of data center power consumption and operational cost.

System-Level Optimization Drives Performance

The performance gains are measured by AgentX, an open-source benchmark designed to replicate the variable and stateful nature of real-world AI agent sessions, including long-context prefill and tool usage. The benchmarks show the GB300 NVL72 itself offers a major leap over its predecessor, delivering up to 15x higher throughput per megawatt than the H200 NVL8 on the DeepSeek V4 Pro model and an 80x gain on larger Mixture-of-Experts (MoE) models like Kimi K3 2.8T. This efficiency translates to up to a 10x reduction in cost per million tokens. According to NVIDIA, these results stem from a co-designed stack including:

  • Optimized MoE serving runtimes like SGLang and TensorRT-LLM.
  • Efficient DeepGEMM-based kernels and mixed-precision formats such as MXFP4.
  • The NVIDIA Dynamo session-aware serving stack for intelligent workload routing.
  • High-bandwidth NVLink fabric connecting 72 GPUs for coordinated rack-scale inference.

This focus on throughput per megawatt directly addresses the critical challenge of scaling AI factories economically. For cloud providers and large enterprises, these efficiency improvements mean they can support a significantly larger volume of interactive agentic AI services within existing power and infrastructure budgets. By demonstrating massive performance-per-watt gains on workloads that are rapidly becoming the industry standard, NVIDIA is reinforcing its market position by solving the pressing operational and financial hurdles associated with deploying next-generation AI at scale.

NVIDIA's latest benchmarks demonstrate a strategic shift from chasing raw performance to optimizing system-level throughput-per-watt, a crucial economic lever for AI factories scaling up stateful, multi-turn agentic workloads.
End of Transmission
Scan All Nodes Access Archive