OpenAI Publishes First Performance Results for Its Jalapeño Custom Inference Chip
OpenAI says Jalapeño, its first custom inference chip, delivered higher peak throughput per kilowatt and lower token latency than the commercial systems in its InferenceX comparison using GPT-OSS 120B, with strong results across additional model families.
What OpenAI announced
On August 25, 2026, OpenAI published the first measured performance results for Jalapeño, its first custom inference chip. The company framed the chip as one layer in a vertically integrated compute strategy spanning data centers, hardware, frontier models, the developer platform and end-user products.
The first benchmark claims
OpenAI says Jalapeño was evaluated on InferenceX, a public inference benchmark, using GPT-OSS 120B. In the comparison reported by OpenAI, the chip delivered more peak throughput per kilowatt and lower token latency than the commercial systems being compared. OpenAI also reported strong results on DeepSeek R1 and Kimi K2, arguing that the efficiency gains are not limited to one model family.
These are OpenAI's own initial measurements, so independent benchmarking across workloads, batch sizes, software stacks and deployment conditions will be important before broader conclusions are drawn.
Why a custom inference chip matters
Inference cost and latency increasingly determine how practical large-scale AI services are. A custom accelerator can be optimized for a provider's serving patterns, memory movement, networking and software stack rather than targeting every possible workload. If the early efficiency results hold across production conditions, Jalapeño could reduce the power and infrastructure required to serve large models while improving response speed.
The broader full-stack strategy
OpenAI's announcement emphasizes co-design: software can make hardware more productive, while hardware tailored to AI workloads can improve model-serving speed and efficiency. The strategic implication is that competition among AI labs is expanding beyond models into chips, data-center architecture, networking and systems software. For developers and enterprises, this may eventually influence inference pricing, latency, model availability and the economics of deploying agentic or high-volume AI applications.
What to watch next
The most useful next evidence will be independent benchmark reproduction, detailed power and utilization data, performance across different model sizes and quantization settings, and information about when and where Jalapeño enters production service. Until then, the August 25 results are best treated as an important first disclosure rather than a complete picture of the chip's production performance.
This article is built from the source material below. Open the originals for full context and the latest updates.