NVIDIA's Vera Rubin NVL72 benchmarks 30x more work per watt and 35x lower token cost for agentic AI. Here is what it really means for agent-fleet economics.
Key takeaways
- On August 24, 2026, NVIDIA published Vera Rubin NVL72 benchmarks showing up to 30x higher throughput per megawatt than GB300 NVL72 on agentic workloads.
- It also reports up to 35x lower cost per million tokens, measured on the SemiAnalysis AgentX benchmark of real agentic coding sessions.
- The gains come from Rubin GPUs, NVLink 6 interconnects with 10x higher packet rates, and NVFP4 4-bit inference via TensorRT LLM and NVIDIA Dynamo.
- The reason efficiency matters: agents grow context, call tools, and spawn sub-agents, so they consume far more tokens than a chat query, making energy and token cost the real bottleneck.
- Van Data Team's recommendation: don't wait on hardware. Model your agent token economics now, and design agents to be efficient and portable so the cost curve works for you.
📖 Read the full guide on Van Data Team → NVIDIA Vera Rubin NVL72: 30x More Work Per Watt
Top comments (0)