Last week at Hot Chips 2026, OpenAI stood on stage and showed benchmarks for its first custom chip. Not a roadmap. Not a "vision." Actual silicon, actual numbers, running actual models.
It's called Jalapeño. Built with Broadcom on TSMC's 3nm-class process, with Samsung supplying the HBM4. And according to OpenAI's own published data, it beats Nvidia's GB200 and GB300 rack systems on the metric that actually matters for a company burning billions of dollars a year on inference: throughput per watt.
If you've spent the last three years assuming Nvidia's position was untouchable, it's time to update that prior.
The numbers
Here's what OpenAI put in front of the room at Hot Chips:
- 1.5x–1.9x higher throughput per kilowatt vs. Nvidia's GB200/GB300 rack systems
- 1.7x–3.6x lower end-to-end latency
- Sustained power draw at or below 550W against a 700W TDP — while Nvidia's comparable dies pull 900–1,150W
- On Kimi-K2.5, Jalapeño hit roughly 700 tokens/sec per user, more than 9x the next best chip tested
- On GPT-OSS, it nearly doubled GB200's peak throughput and beat GB200's concurrency-1 number by more than 50x
SemiAnalysis, who got early access to run the numbers, didn't hedge: Jalapeño is "beating every Nvidia, AMD, and Google chip we have been able to test." That's not marketing copy from OpenAI's PR team — that's the most plugged-in independent chip analysis shop in the industry saying the quiet part out loud.
And this is a first-generation chip. Dylan Patel's line is the one worth sitting with: "Usually first generation chips aren't competitive." OpenAI's isn't just competitive — it's winning most of the categories that end up on an infra bill.
Under the hood
Each Jalapeño package pairs a reticle-sized compute die with six HBM4 stacks — 216 GiB at 15.4 TB/s of bandwidth, with HBM4 pin speeds running at 10Gbps versus Nvidia Rubin's 9.6Gbps. The B0 stepping delivers 13.4 PFLOPs MXFP4 per die, against Rubin's 17.5 PFLOPs dense NVFP4 — so on raw compute, Nvidia still wins. Jalapeño wins on efficiency and latency instead, which is the number that shows up on your cloud bill, not the number that shows up on a spec sheet.
The architecture choices tell you OpenAI actually understands its own workload, not just "we can build a chip too":
- Weight-stationary systolic array, closer to a TPU than a GPU
- Out-of-order cores with L1 cache — a genuine departure from typical accelerator design, aimed at flexibility over raw throughput
- No prefill-decode disaggregation. Everyone else in the industry is splitting prefill and decode across separate hardware pools. OpenAI didn't, because it preserves KV cache locality and keeps utilization flexible as workload mixes shift. That's a bet made by people who run the workload at scale, not people copying a paper.
They also wrote a whole software stack to go with it: Gluon, a kernel language built on Triton, and "Teacup," their internal serving engine. When they didn't have MLA kernels for DeepSeek, they didn't write them by hand — they had Codex generate them. Sixteen months from design to tape-out. Two-week turnarounds on 2x throughput gains. That's not typical chip company velocity. That's a software company that decided hardware is just another compiler target.
The caveats, because there always are some
Before you short NVDA on this article: the comparison to Blackwell is, by SemiAnalysis's own admission, "somewhat incomplete and unfair." Blackwell doesn't use HBM4. The apples-to-apples comparison is Vera Rubin, and there Jalapeño is roughly at parity on cost-per-token, not a clear winner. The benchmarks were run on 8k-input/1k-output workloads — not the long-context, multi-turn agentic loads that are increasingly what production traffic actually looks like. No AgentX numbers were published. The numbers come from OpenAI. Production ramp doesn't start meaningfully until 2027.
So no, Nvidia is not dying. But that's not the point.
Why this actually matters
CUDA is still the moat. TensorRT, cuDNN, the entire ecosystem of tooling every ML team has built their workflow around for a decade — none of that disappears because one customer shipped a chip. If you're a startup fine-tuning open models, your stack doesn't change tomorrow.
What changes is Nvidia's pricing power, and pricing power is the whole business model. For years the pitch to hyperscalers was: there is no alternative, so pay what we ask. OpenAI just spent 16 months and a ton of engineering headcount proving there is an alternative — for their own workload, on their own money, at their own scale of $65B+ annualized revenue. Google's been running TPUs for years. Amazon has Trainium. Now the company that arguably drove the entire GPU shortage in the first place has its own chip beating Nvidia's flagship on the number that shows up in the finance meeting.
Every one of Nvidia's biggest customers is quietly becoming a competitor. That's not a moat cracking today. That's a moat with a very visible waterline dropping, quarter over quarter, and everyone in the room can see it.
If you build on top of any of these models, here's the actual takeaway: don't architect around today's GPU pricing assuming it's permanent. The chip layer under LLM inference is about to get genuinely competitive for the first time since the current AI boom started. Cheaper inference is coming — the only open question is whether the savings get passed down to you, or absorbed as margin. Based on how this industry has behaved so far, I wouldn't bet on the former.
Top comments (0)