DEV Community

gentic news
gentic news

Posted on Originally published at gentic.news

OpenAI Jalapeño Chip Beats Nvidia Blackwell on InferenceX

OpenAI's Jalapeño chip beat Nvidia Blackwell on InferenceX at Hot Chips 2026, with more tokens per watt. Volume deployment slips to 2027, raising questions about the competitive window.

OpenAI's Jalapeño inference chip beat Nvidia's Blackwell on SemiAnalysis' InferenceX benchmark at Hot Chips 2026. Richard Ho, OpenAI's head of hardware, called the results a "very, very significant performance advance over state of the art."

Key facts

  • Jalapeño beats Nvidia Blackwell on SemiAnalysis InferenceX benchmark
  • More tokens per user and throughput per kilowatt than state-of-the-art
  • Small deployment late 2026, significant volume in 2027
  • Co-developed with Broadcom, announced October 2025
  • Designed to minimize prefill and communication phase delays

OpenAI's custom inference chip Jalapeño outperformed Nvidia's Blackwell system on SemiAnalysis' InferenceX benchmark, registering more tokens per user and more throughput per kilowatt, the company disclosed at Hot Chips on Tuesday. The comparison is notable, but the competitive window is narrow: Ho estimated Jalapeño would deploy "in very small volumes" at the end of 2026, with meaningful scale only in 2027 — by which point Nvidia's next-generation parts will likely be shipping. According to TechCrunch

The prefill and KV cache angle

The benchmark win is less about raw silicon and more about where inference bottlenecks actually live. Jalapeño was designed with Broadcom to minimize delays during prefill and communication phases, and to keep the KV cache local. "We designed Jalapeño to minimize data movement and communication delays," OpenAI said in a blog post, "so that model state, including the KV cache used while generating a response, can be explicitly placed and kept local while the system activates the right combination of compute, memory, and networking for each inference phase." This is a direct response to the agentic inference workloads that have made KV cache management the central problem of the serving stack — a theme SemiAnalysis has pushed with its open-sourced $3M AgentX-InferenceXv3 dataset released just a day earlier.

The timing problem

The more uncomfortable question is whether beating today's Blackwell matters when the comparison target will be obsolete by the time Jalapeño ships at volume. Ho acknowledged the deployment timeline openly. OpenAI is also cutting API prices aggressively — GPT-5.6 Sol prices dropped 20-33% this week — which makes per-watt throughput a strategic lever, not just a technical one. The full-stack approach, with models, chips, and memory developed in concert, is the real structural advantage: it lets OpenAI attack inference phases that generic accelerators treat as uniform.

The company did not disclose absolute token counts, power draw, or die size, which limits the extent to which the InferenceX result can be independently verified. The company's blog post presents the results as a comparison against "currently available" state-of-the-art parts, a phrasing that leaves room for interpretation about what comes next.

Key Takeaways

  • OpenAI's Jalapeño chip beat Nvidia Blackwell on InferenceX at Hot Chips 2026, with more tokens per watt.
  • Volume deployment slips to 2027, raising questions about the competitive window.

What to watch

Watch for Nvidia's next-generation inference parts and whether Jalapeño's InferenceX advantage holds against them. Also track OpenAI's per-watt cost data in Q1 2027 as volume deployment begins, and whether the GPT-5.6 Sol price cuts reflect Jalapeño's efficiency or competitive pressure ahead of Anthropic's IPO.

OpenAI’s Jalapeño chip


Source: techcrunch.com

[Updated 25 Aug via the_decoder]

The chip's advantage extends beyond Blackwell: SemiAnalysis CEO Dylan Patel said Jalapeño also beats Nvidia's upcoming Rubin, noting "usually first generation chips aren't competitive, but OpenAI is beating Nvidia Blackwell and even Rubin." [per The Decoder] Specific figures from SemiAnalysis show Jalapeño delivers 1.5 to 1.9 times more AI work per watt at peak throughput, 1.7 to 3.6 times lower end-to-end latency, and 2.1 to 4.1 times higher performance on highly interactive workloads. [per Next Big Future]


Originally published on gentic.news

Top comments (0)