“AI Disruption” Publication 10,500 Subscriptions 20% Discount Offer Link.
Making chips is hard; making a competitive first-generation product is even harder. The answer OpenAI has delivered this time is, at least when it comes to large-model inference, already quite extraordinary.
Today, OpenAI officially released the first batch of measured performance data for its first custom inference chip, Jalapeño (Mexican chili pepper).
On the public benchmark InferenceX based on its own open-source model GPT-OSS 120B, Jalapeño outperforms existing hardware systems in both peak throughput per kilowatt and token latency.
It’s worth noting that for most chips, throughput and latency are often a trade-off where you can only choose one. Beyond the clear optimization for OpenAI’s large models, the chip also performs better on the 670B DeepSeek R1 and the 1T Kimi K2.5, with the competitors being the strongest commercial AI systems on the market, such as NVIDIA’s GB200/GB300.
Altman commented bluntly on X: “We built a chip that’s ridiculously fast.”
From data centers to phones, AI chips are becoming a contested territory for all tech giants.




