Editorial illustration for OpenAI's Jalapeño Chip Boasts High Efficiency, Low Latency in Benchmarks
OpenAI's Jalapeño Chip Beats Nvidia in Efficiency Test
OpenAI's Jalapeño Chip Boasts High Efficiency, Low Latency in Benchmarks
OpenAI put numbers behind its custom silicon on Tuesday, presenting the first benchmark results for Jalapeño at the Hot Chips conference. The chip, built with Broadcom and first announced last October, ran on SemiAnalysis's InferenceX benchmark against a current Nvidia Blackwell system. Jalapeño came out ahead on two measures that matter most for running AI models at scale: tokens delivered per user and throughput per kilowatt of power consumed.
Richard Ho, OpenAI's head of hardware, walked reporters through the results on a press call. The company designed Jalapeño as part of a multigenerational platform where chips, memory, models, and products get built together rather than in isolation, and OpenAI says its own models helped shape the chip's development. That full-stack approach let engineers target specific pain points in the inference pipeline, particularly the prefill and communication phases that tend to slow systems down when serving AI workloads at volume.
Full deployment is still a ways off. Ho said Jalapeño will ship in small volumes by the end of 2026, with wider rollout planned for 2027, meaning today's Blackwell comparison could look different by the time the chip actually reaches customers.
Tested on SemiAnalysis’ InferenceX benchmark, Jalapeño registered both more tokens per user and more throughput per kilowatt than the currently available state-of-the-art inference processors.
Why this matters Every inference benchmark OpenAI publishes is also a negotiating chip, literally, against Nvidia's pricing and roadmap. The Hot Chips numbers on Jalapeño look real: more tokens per user, more throughput per kilowatt than a current Blackwell system, per SemiAnalysis' InferenceX test. But the comparison is against hardware Nvidia already shipped, not whatever it fields when Jalapeño actually reaches scale. That gap matters more than Richard Ho's framing suggests.
For developers and founders building on OpenAI's API, the pitch is straightforward: cheaper tokens, faster responses, once Jalapeño is in production. For researchers watching the custom-silicon race, this is one data point in a field where Google, Amazon and Microsoft are running the same play. The number worth tracking isn't the benchmark OpenAI chose to release. It's whether Jalapeño's advantage still holds once it's compared against whatever Nvidia has shipped by the time deployment actually happens, not the chip sitting in production today.
Common Questions Answered
How does OpenAI's Jalapeño chip compare to Nvidia's Blackwell system in the SemiAnalysis InferenceX benchmark?
According to the benchmark results presented at the Hot Chips conference, Jalapeño outperformed the current Nvidia Blackwell system on two critical measures: tokens delivered per user and throughput per kilowatt of power consumed. These metrics are particularly important for running AI models efficiently at scale, making Jalapeño's superior performance in these areas a significant achievement for OpenAI's custom silicon.
What companies collaborated to develop the Jalapeño chip?
OpenAI built the Jalapeño chip in partnership with Broadcom, which was first announced in October of the previous year. The collaboration between these two companies resulted in the custom silicon that was benchmarked and presented at the Hot Chips conference.
Why is the Jalapeño chip's efficiency measured in tokens per user and throughput per kilowatt?
These two metrics matter most for running AI models at scale because they directly measure both user experience (tokens per user) and operational cost-effectiveness (throughput per kilowatt). Together, they demonstrate that Jalapeño can deliver faster inference performance while consuming less power, which are critical factors for large-scale AI deployment.
What is the significance of OpenAI publishing Jalapeño benchmark results against current Nvidia hardware?
Publishing these benchmark results serves as a negotiating tool against Nvidia's pricing and product roadmap, as every inference benchmark OpenAI releases has competitive implications. However, the comparison is against hardware Nvidia has already shipped rather than future generations, which means the performance gap between Jalapeño and Nvidia's next-generation offerings when Jalapeño reaches scale could differ significantly from current benchmark results.
Further Reading
- OpenAI’s Jalapeño chip is built for fast inference at scale, benchmarks show - TechCrunch
- OpenAI's upcoming Jalapeño chip looks like it'll be an inference beast - The Register
- OpenAI unveils custom chip it designed with Broadcom to boost its AI infrastructure - Reuters
- OpenAI and Broadcom unveil LLM-optimized inference chip - OpenAI
- OpenAI unveils its first custom chip, built by Broadcom - TechCrunch