AI Disruption

AI Disruption

OpenAI's "Pepper Chip" Beats NVIDIA in First Test

OpenAI's first inference chip Jalapeño beats NVIDIA in throughput and latency, with AI-driven design reshaping the future of AI hardware and CUDA's dominance.

Meng Li's avatar
Meng Li
Aug 26, 2026
∙ Paid

“AI Disruption” Publication 10,500 Subscriptions 20% Discount Offer Link.


Making chips is hard; making a competitive first-generation product is even harder. The answer OpenAI has delivered this time is, at least when it comes to large-model inference, already quite extraordinary.

Image

Today, OpenAI officially released the first batch of measured performance data for its first custom inference chip, Jalapeño (Mexican chili pepper).

On the public benchmark InferenceX based on its own open-source model GPT-OSS 120B, Jalapeño outperforms existing hardware systems in both peak throughput per kilowatt and token latency.
It’s worth noting that for most chips, throughput and latency are often a trade-off where you can only choose one. Beyond the clear optimization for OpenAI’s large models, the chip also performs better on the 670B DeepSeek R1 and the 1T Kimi K2.5, with the competitors being the strongest commercial AI systems on the market, such as NVIDIA’s GB200/GB300.
Altman commented bluntly on X: “We built a chip that’s ridiculously fast.”

From data centers to phones, AI chips are becoming a contested territory for all tech giants.

User's avatar

Continue reading this post for free, courtesy of Meng Li.

Or purchase a paid subscription.
© 2026 Meng Li · Privacy ∙ Terms ∙ Collection notice
Start your SubstackGet the app
Substack is the home for great culture