DEV Community

pablo padlo
pablo padlo

Posted on Originally published at x.com

What an Hour of NVIDIA's B300 Actually Costs in 2026

#ai

What an Hour of NVIDIA's B300 Actually Costs in 2026

Renting NVIDIA's Blackwell Ultra B300 — the GPU the 2026 AI buildout is priced against — costs between $6.94 and $8.71 per GPU-hour on demand. Spot capacity trades at $3.75, and the on-demand price has climbed about 83% in six months. This is what a single hour of the compute behind every frontier model actually costs.

The spread between providers is the first surprise

The same chip sells for three different prices depending on who is renting out the hour. RunPod's community tier lists a single B300 at $6.94 per hour; its secure tier at $7.89. Verda in Finland lists $7.50 with spot capacity at $3.75. Nebius sits at $7.85, with spot at $4.30. On eight-GPU nodes the arithmetic improves: Excesssupply advertises $36 an hour for eight B300s, which is $4.50 per GPU.

Third-party trackers give the medians buyers actually benchmark against. Cloud-gpus.com (11 September 2026) reports $8.71 on demand, $3.75 on spot and $7.24 on reserved capacity across five providers. Gambit.zone's 10 August snapshot puts the B300 median at $7.94 across seven quotes, with a range from $7.17 to $17.80. Gputable.dev tracks a cheapest-listed floor of $6.60.

The card is the cheap part

A single B300 is estimated at roughly $50,000. A DGX B300 with eight of them lists near $350,000. A full GB300 NVL72 rack — 72 B300s, 36 Grace CPUs, 37 terabytes of fast memory — is estimated between $3M and $6.5M depending on how the measurement is done.

That last figure is where the market is heading. Microsoft has put the first at-scale GB300 NVL72 cluster into production for OpenAI workloads, and Azure's ND GB300 v6 VMs are being sold as the standard infrastructure for multitrillion-parameter models.

Why the price is going up

The on-demand number is the visible artifact of a supply constraint. Across the pricing boards, on-demand B300 capacity went from roughly $4.5 an hour to $9.16 in about six months — a rise of 83%. Reserved contracts are currently around 2.7 times cheaper than the on-demand spike, a wider spread than at any point in the H100 era.

That gap is the honest signal. Both vendors and end users report supply constraints, and the width of the reserved discount tells you how much of the on-demand price is scarcity rather than cost.

What the silicon actually buys

Each B300 carries 288 GB of HBM3e, 8 TB/s of memory bandwidth and 18 PFLOPS of FP4 inference performance. The precision and the bandwidth are the point: long-context agentic inference is bottlenecked on exactly that combination, because a large KV cache has to stay in the fastest memory tier rather than being offloaded.

NVIDIA frames the workload as test-time scaling — the third scaling dimension after pretraining and post-training. The company's own documentation notes that reasoning at inference can demand up to 100 times the compute of one-shot generation, which is the demand curve the B300 was designed around.

What to watch

NVIDIA's headline figures are 35 times cheaper tokens and 50 times the throughput per megawatt versus Hopper, based on SemiAnalysis InferenceX benchmarks. On DeepSeek-R1, Blackwell Ultra systems reached 2.5 million tokens per second in MLPerf Inference v6.0, and NVIDIA quotes $0.24 per million tokens as an inference reference.

Those are vendor and partner numbers. The number that decides an actual bill is utilization: a node charges for the hour whether the silicon is working or waiting. That is not on any datasheet, and it is the one worth watching.

Watch the rack (video):

https://x.com/pgol80/status/2102364543091376230

Sources:

Top comments (0)