OpenAI presents Jalapeño inference benchmarks
OpenAI has presented initial benchmarks for its first in-house AI inference chip, called “Jalapeño,” at the Hot Chips conference. The company says the accelerator surpasses Nvidia’s Blackwell and Rubin systems in throughput per watt and token latency.
Jalapeño is designed only for inference: it runs AI models but does not train them. OpenAI also says it is a general-purpose large language model accelerator rather than hardware optimized specifically for OpenAI models.
Across three tested models, OpenAI reports 1.5x to 1.9x higher peak AI work per watt. End-to-end latency was reportedly 1.7x to 3.6x lower than on the best commercially available systems, while interactive workload performance was 2.1x to 4.1x higher.

Matthias BastianAug 25, 2026

Benchmark results
The results came from SemiAnalysis’s public InferenceX benchmark. OpenAI supplied the figures, and SemiAnalysis verified some test runs on-site. The evaluated models were GPT-OSS 120B, Deepseek R1 670B, and Kimi K2.5 1T.
On GPT-OSS, Jalapeño reached approximately 1,400 tokens per second per user. On Deepseek R1, it exceeded 700 tokens per second with one concurrent request.
At equivalent decoding speeds, OpenAI reports 54x to 104x greater token throughput per kilowatt than the best available accelerator, depending on the model.
Jalapeño achieved these results without multi-token prediction or speculative decoding. Some competing systems used those techniques, leaving potential room for Jalapeño to improve. SemiAnalysis described the chip as outperforming every other chip in its headline performance-per-watt comparison. CEO Dylan Patel said it was unusual for a first-generation chip to compete with Nvidia’s Blackwell and Rubin.
Comparison with Vera Rubin
SemiAnalysis argues that Nvidia’s Vera Rubin platform is the more appropriate comparison because both systems use HBM4 memory. Jalapeño still produced more output tokens per megawatt than Vera Rubin, despite Nvidia’s accelerator using multi-token prediction. However, the two systems were roughly equal in total cost of ownership per token.
The comparison has limitations. Nvidia and AMD have published results for larger models, including Deepseek V4 Pro and Kimi K3, that have not yet been tested on Jalapeño. Rubin systems are already shipping to customers, while Jalapeño reportedly remains at the engineering-sample stage.
Development with Broadcom
OpenAI developed Jalapeño with Broadcom. Design work began in mid-2024, and the final design entered fabrication in November 2025. Although the overall process took about 16 months, OpenAI says only nine months elapsed between the first chip design and the completed blueprint sent to the factory.
OpenAI used its own AI models during development. Older models assisted with chip design, while newer generations helped accelerate programming and optimization.
Part of a broader compute strategy
SemiAnalysis said the rapid development could challenge Nvidia’s long-discussed CUDA advantage. OpenAI CFO Sarah Friar described Jalapeño as part of an integrated strategy spanning data centers, chips, models, the developer platform, products, and devices.
Friar said the chip complements, rather than replaces, OpenAI’s partnerships with Nvidia, AMD, AWS, Cerebras, CoreWeave, and other providers. Several of those companies are both partners and competitors: Nvidia, AMD, and AWS are investors or compute partners, while each is also developing its own AI hardware.