AWS and NVIDIA announced plans to deploy 2 million additional NVIDIA GPUs across AWS global infrastructure during 2027 and 2028, expanding a partnership that already committed more than a million starting in 2026. The release also commits 100,000 GPUs to AWS secure infrastructure for US federal and national-security workloads, brings NVIDIA's Vera CPUs to AWS, and extends NVIDIA's NVLink Fusion interconnect to Amazon's own Trainium chips.
Key facts
- 2 million additional NVIDIA Blackwell Ultra, Rubin and Rubin Ultra GPUs across AWS Global Infrastructure in 2027-2028.
- On top of the more than 1 million announced at NVIDIA GTC in March 2026; the August release says demand exceeded those expectations.
- 100,000 GPUs planned on AWS secure infrastructure for the US Government, supporting workloads at Impact Level 6 and above.
- Announced August 26, 2026. Primary source: AWS and NVIDIA press release.
The hook. Jensen Huang, NVIDIA's founder and CEO, put the demand picture in one sentence: "NVIDIA and AWS have built one of the great growth engines of the AI era, and demand is running ahead of every forecast."
Background. Cloud providers announce capacity in units that are hard to hold in your head. The useful frame is that this lands in the same month as reporting that a small number of frontier labs have already contracted a large share of next year's available compute -- we covered that in two labs took about thirty percent of this year's new compute. Additional supply and concentrated demand are the two halves of the same question: whether anyone outside the biggest labs can get chips.
What was announced. The GPU number is the headline, but the co-engineering items are what change the architecture. NVIDIA Vera CPUs are coming to AWS as an option for agentic workloads needing heavy CPU compute alongside accelerators. NVLink Fusion, NVIDIA's high-speed chip interconnect, is being extended to work with NVHBM custom memory on Amazon's Annapurna Labs Trainium silicon, which the release says lets Trainium and NVIDIA GPUs sit inside a common rack-scale architecture. New G7 instances built on RTX PRO 4500 Blackwell Server Edition GPUs claim 4.6 times the AI inference performance and 2.1 times the graphics performance of the previous G6 generation, with AWS the first major cloud to offer them. NVIDIA Spectrum networking is being tuned for large-scale training across GPU clusters, and NVIDIA's open Nemotron models are coming to AWS.
How it works. A modern AI data centre is less a pile of chips than a memory system with compute attached. The bottleneck for both training and serving is usually how fast data moves between accelerators, not how fast any single accelerator calculates -- the same physics behind why LLM inference is memory-bound. NVLink Fusion is the fabric that ties accelerators together at near-local speed. Extending it to Trainium means Amazon's in-house chips and NVIDIA's can share that fabric instead of living in separate racks. Practically: AWS gets to sell whichever silicon a customer wants without splitting its data-centre design in two.
Why it matters. Matt Garman, CEO of AWS, framed the strategy as choice: "Customers want the freedom to choose the best tools for their AI workloads, and they want confidence that everything works seamlessly together. That's why we've invested deeply with NVIDIA to make AWS the best place to run NVIDIA AI technologies." The federal commitment is the quieter item. AI factories for the US Government at Impact Level 6 and above puts AWS and NVIDIA jointly at the centre of classified national-security AI workloads, a market with different procurement rules and far less price sensitivity than commercial cloud.
The honest caveat. Two. First, these are plans across a two-year window, not deployed capacity -- Blackwell Ultra, Rubin and Rubin Ultra span multiple hardware generations, and announced cloud capacity has a long history of slipping. Second, the framing that carried this story on aggregators was that Amazon "tripled" its Nvidia order, and that is not what either release says. March committed to more than a million starting in 2026; August adds two million more in 2027-2028. It is a large expansion in a later window, arithmetically distinct from multiplying an existing order, and the difference matters if you are trying to model when chips actually become available.
Originally published on Ground Truth, where every claim is checked against the primary source.
Top comments (0)