Key Takeaways
- OpenAI’s Jalapeño ASIC, co-designed with Broadcom and fabricated by TSMC, went from RTL design to first silicon in nine months, with AI models writing more than half of the core.
- On OpenAI’s own benchmarks at Hot Chips 2026, Jalapeño delivered 1.5 to 1.9 times higher throughput per kilowatt and 1.7 to 3.6 times lower latency than Nvidia’s GB200 and GB300 on specific inference workloads; no independent verification has been published.
- OpenAI is now running a multi-vendor supply chain, Cerebras, AWS Trainium, AMD and its own silicon, putting direct pressure on inference margins that Nvidia’s $1 trillion Blackwell and Vera Rubin forecast has not yet absorbed. Jalapeño went from architectural concept to running OpenAI‘s Codex model in nine months, with AI models writing more than half the chip’s core in the XLS hardware language. The performance numbers OpenAI presented at Hot Chips 2026 explain why the company wanted its own silicon rather than buying more Nvidia.
Jalapeño’s Performance Profile
Jalapeño is an ASIC built specifically for LLM inference, not general-purpose GPU workloads. That focus shows in the numbers OpenAI presented at Hot Chips 2026: against Nvidia‘s GB200 and GB300 on specific workloads, Jalapeño delivered 1.5 to 1.9 times higher peak throughput per kilowatt and 1.7 to 3.6 times lower end-to-end latency, according to OpenAI’s own testing. No independent verification of those figures has been published.
The sharpest comparison came on a 1-trillion-parameter Kimi K2.5 run: Jalapeño drawing 700W against the GB300’s 1,400W, with 1.5 times higher peak mixed tokens per second per kilowatt and 3.4 times lower end-to-end latency. OpenAI frames its metrics around “time to last token” for user experience and “tokens per joule” for efficiency, framing that positions Jalapeño as an inference platform rather than a raw accelerator.
Nine Months from RTL to Silicon
OpenAI moved from initial RTL design to tapeout in roughly nine months, with AI models actively searching for power, performance and area improvements throughout. More than half of the core was written using the XLS hardware language. First silicon arrived in May 2026; Codex was running on it the same month. Greg Brockman, OpenAI’s president, said the chip should cut operational costs significantly, according to reports.
The development speed matters beyond the headline. Custom silicon programs at established semiconductor companies typically run two to three years from architecture to first silicon. Nine months, with AI-assisted RTL generation covering more than half the core, is the more consequential data point buried inside OpenAI’s inference cost problem: the company is not just building a chip, it is proving a faster path to custom hardware.
Nvidia’s $1 Trillion Forecast
Nvidia has forecast total sales reaching $1 trillion across its Blackwell and Vera Rubin architectures for 2026 and 2027 combined, according to the company’s projections. Those numbers suggest custom silicon from hyperscalers has not yet dented Nvidia’s revenue trajectory in any meaningful way. Training remains almost entirely GPU-dominated, and Nvidia’s next-generation architectures are designed to defend its inference position too.
The pressure from custom ASICs is real but concentrated. Jalapeño targets a narrow efficiency band, high-volume, specific-workload inference, where a purpose-built chip can undercut a general-purpose GPU on watts per token. Outside that band, Nvidia’s programmability and software stack still win. The question is how wide that band gets as inference spending scales past training.
A Wider Hardware Shift
Google has been building Tensor Processing Units since 2015. Amazon’s Trainium launched in December 2020. Meta and Microsoft have their own silicon programs. What OpenAI adds is a different kind of pressure: one of Nvidia’s largest customers is now designing around it at scale.
OpenAI is running a multi-vendor supply chain in parallel: Cerebras for its GPT-5.6 Sol model, AWS Trainium through a multibillion-dollar Amazon deal, and a separate agreement with AMD. The competitive pressure lands squarely on inference margins. If Jalapeño performs as OpenAI’s benchmarks suggest, the economics of large-scale LLM deployment shift, and given that enterprise appetite for non-Nvidia silicon is already measurable, OpenAI is not the only buyer watching.
Originally published at https://autonainews.com/openais-jalapeno-chip-claims-inference-efficiency-over-nvidia/
Top comments (0)