OpenAI published the first benchmarks for its custom Jalapeño inference chip. Here is what the numbers say, and what they mean if you build on OpenAI.
Key takeaways
- On August 25, 2026, OpenAI published the first benchmarks for its Jalapeño inference chip, run on the public SemiAnalysis InferenceX benchmark.
- OpenAI reports 1.5 to 1.9 times more work per watt at peak throughput and 1.7 to 3.6 times lower end-to-end latency than the comparison systems, with up to 2.1 to 4.1 times higher performance on interactive workloads.
- The gains held across three open models, GPT-OSS 120B, DeepSeek R1 670B, and Kimi K2.5 1T, suggesting a chip-level advantage rather than a model-specific trick.
- Two honest caveats: these are OpenAI's own numbers pending independent testing, and the chip ships in very small volumes at the end of 2026, with real deployment in 2027.
- Van Data Team's recommendation: treat this as a strategic signal about first-party silicon and vendor choice, not a reason to change what you build today.
📖 Read the full guide on Van Data Team → OpenAI's Jalapeño Chip: First Inference Benchmarks
Top comments (0)