OpenAI has released the first performance results for Jalapeño, its custom-built AI inference chip, showing notable improvements in speed, latency, and power efficiency.
Inference is the process of running a trained AI model to generate a response. As AI services become more widely used, faster and more efficient inference can help reduce operating costs while improving response times for users.
According to OpenAI, Jalapeño delivered 1.5 to 1.9 times more AI work per watt at peak throughput and 1.7 to 3.6 times lower end-to-end latency than the comparison systems used in testing.
The chip was evaluated on three large models: GPT-OSS 120B, DeepSeek R1 670B, and Kimi K2.5 1T. OpenAI says Jalapeño performed strongly across all three, suggesting the architecture can support models beyond its own ecosystem.
One of Jalapeño’s key design goals is to reduce the delays caused by moving data between compute, memory, and networking resources. OpenAI says the chip and its surrounding system were designed together specifically for modern language-model workloads, including increasingly interactive AI agents.
AI also played a role in Jalapeño’s development. OpenAI says its models helped engineers speed up parts of the chip design, verification, and software optimization process. In selected GPT-OSS components, AI-generated implementations reportedly ran 1.5 to 1.8 times faster than existing human-written versions.
The early results are promising, but they should still be viewed as preliminary. SemiAnalysis says it verified benchmark runs in OpenAI’s lab, while the published performance figures themselves were supplied by OpenAI. Broader independent testing and real-world deployment will provide a clearer picture of how Jalapeño performs at scale.
OpenAI plans to begin deploying Jalapeño in its own infrastructure by the end of 2026. The company says second- and third-generation versions are already in development, while NVIDIA and other partner accelerators will continue to remain part of its broader compute strategy.
If the early results hold up in production, Jalapeño could help OpenAI serve more AI workloads with lower latency and better energy efficiency, while giving the company greater control over a critical part of its infrastructure.

Leave a Reply