OpenAI Publishes First Jalapeno Chip Performance Results
OpenAI published the first measured performance results for Jalapeno, its first custom inference chip, on August 25, 2026. Everything below comes from that first party post.
Transcript
OpenAI just published the first measured results for Jalapeno, its first custom inference chip. These are OpenAI's own tests.
OpenAI ran Jalapeno on InferenceX, a public benchmark from SemiAnalysis, against leading commercial systems. Across three models it reports up to one point nine times more work per watt.
Jalapeno is rated at seven hundred watts, but OpenAI measured draw at or below five hundred fifty. On the largest model tested, latency was three point four times lower.
OpenAI says its own models helped take Jalapeno from design to tapeout in nine months. Deployment inside OpenAI begins by year end. It will keep using NVIDIA accelerators too.
Inboxsmith builds AI receptionists that answer calls for small businesses. Please like and subscribe for more news.
Sources
Every claim in this video comes from the top ranking coverage of this topic. The claims and where each one came from:
- Since announcing Jalapeno, OpenAI's first custom inference chip, we have been testing the chip and the system built around it.(OpenAI's official announcement, Jalapeno's first results show industry-leading speed and efficiency in AI inference, published August 25, 2026)
- To understand how Jalapeno performs in practice, we tested it on InferenceX, a public benchmark from SemiAnalysis that measures the full process of serving an AI request. We compared Jalapeno with leading commercially available AI systems across the tested operating range, from high-throughput serving to highly interactive, low-latency use.(OpenAI's official announcement, Jalapeno's first results show industry-leading speed and efficiency in AI inference, published August 25, 2026)
- Jalapeno's performance extends across GPT-OSS 120B, DeepSeek R1, and Kimi K2.5 1T, showing that the architecture works across models developed both inside and outside OpenAI.(OpenAI's official announcement, Jalapeno's first results show industry-leading speed and efficiency in AI inference, published August 25, 2026)
- Across all three, Jalapeno delivered 1.5 to 1.9 times more AI work per watt at peak throughput and 1.7 to 3.6 times lower end-to-end latency than the comparison systems.(OpenAI's official announcement, Jalapeno's first results show industry-leading speed and efficiency in AI inference, published August 25, 2026)
- Jalapeno is rated at 700 watts, although its measured sustained power remained at or below 550 watts on the workloads tested.(OpenAI's official announcement, Jalapeno's first results show industry-leading speed and efficiency in AI inference, published August 25, 2026)
- On Kimi, the largest public model we tested, it delivered approximately 1.5 times higher peak performance per watt and 3.4 times lower end-to-end latency than the comparison system.(OpenAI's official announcement, Jalapeno's first results show industry-leading speed and efficiency in AI inference, published August 25, 2026)
- AI played a direct role in Jalapeno's development, enabling the team to move from initial design to tapeout in nine months by exploring implementations, shortening design, measurement, and verification loops, and continuously iterating on model workloads.(OpenAI's official announcement, Jalapeno's first results show industry-leading speed and efficiency in AI inference, published August 25, 2026)
- We plan to begin deploying Jalapeno within OpenAI's compute infrastructure by the end of the year. We will continue to widely deploy accelerators from NVIDIA and other partners for both training and inference workloads.(OpenAI's official announcement, Jalapeno's first results show industry-leading speed and efficiency in AI inference, published August 25, 2026)
