DEV Community

Tran Tien Van
Tran Tien Van

Posted on Originally published at vandatateam.com

NVIDIA Vera Rubin NVL72: 30x More Work Per Watt

NVIDIA's Vera Rubin NVL72 benchmarks 30x more work per watt and 35x lower token cost for agentic AI. Here is what it really means for agent-fleet economics.

Key takeaways

  • On August 24, 2026, NVIDIA published Vera Rubin NVL72 benchmarks showing up to 30x higher throughput per megawatt than GB300 NVL72 on agentic workloads.
  • It also reports up to 35x lower cost per million tokens, measured on the SemiAnalysis AgentX benchmark of real agentic coding sessions.
  • The gains come from Rubin GPUs, NVLink 6 interconnects with 10x higher packet rates, and NVFP4 4-bit inference via TensorRT LLM and NVIDIA Dynamo.
  • The reason efficiency matters: agents grow context, call tools, and spawn sub-agents, so they consume far more tokens than a chat query, making energy and token cost the real bottleneck.
  • Van Data Team's recommendation: don't wait on hardware. Model your agent token economics now, and design agents to be efficient and portable so the cost curve works for you.

📖 Read the full guide on Van Data Team → NVIDIA Vera Rubin NVL72: 30x More Work Per Watt

Top comments (0)