OpenAI Unveils Test Results for First Custom Inference Chip "Jalapeño," Leading in Energy Efficiency and Latency
On August 25, OpenAI released preliminary test results for its first custom inference chip, Jalapeño. The data shows that on models such as GPT-OSS 120B, DeepSeek R1 670B,…
On August 25, OpenAI released preliminary test results for its first custom inference chip, Jalapeño. The data shows that on models such as GPT-OSS 120B, DeepSeek R1 670B, and Kimi K2.5 1T, the chip achieves a 1.5x to 1.9x improvement in peak performance per watt compared to baseline systems, and reduces end-to-end latency by 1.7x to 3.6x; under high-interaction workloads, performance gains reach 2.1x to 4.1x.
OpenAI stated that Jalapeño leverages co-design across chip, memory, networking, software, and rack-level systems to reduce data movement and communication latency, balancing high throughput with low latency. The company plans to deploy Jalapeño into its computing infrastructure by the end of this year, and noted that its second-generation chip is already in deep development, with a third-generation chip also in the planning stages.
[TechFlow]
Original: https://www.techflowpost.com/newsletter/detail_133484.html
insigtX content is informational and educational, not investment advice.