OpenAI has published the first benchmark results for Jalapeño, its custom inference chip, and the numbers are striking: on SemiAnalysis' InferenceX benchmark, the new silicon registered both more tokens per user and more throughput per kilowatt than the currently available state-of-the-art inference processors — a comparison that includes Nvidia's Blackwell systems. The results were presented Tuesday at the Hot Chips conference in Palo Alto, alongside a detailed technical look at how the chip was built. For readers tracking the compute race that underpins every AI model release, it is one of the most consequential hardware datapoints of the year.

"The bottom line is that the results show a very, very significant performance advance over state of the art," said Richard Ho, OpenAI's head of hardware, in a press call reported by TechCrunch. "Jalapeño can serve more AI work per unit of power, while also returning responses more quickly. It's very efficient to serve a lot of customers, but it can also be very low latency."

A Chip Built for the Full Inference Pipeline

Jalapeño was unveiled in June 2026 as OpenAI's first custom silicon, developed in close collaboration with Broadcom, with OpenAI's own AI models assisting in the chip's development process. The partnership itself dates back to October 2025, when OpenAI laid out plans for a multigenerational custom accelerator program alongside the networking giant.

What makes Jalapeño different from a general-purpose GPU, according to OpenAI, is that it was designed in concert with the models it will serve. Because the company controls the full stack — AI products, models, chips, and memory — its engineers were able to attack specific phases of the inference process that often cause friction.

In particular, Jalapeño is designed to minimize delays during the prefill and communication phases of processing, which OpenAI says frequently act as bottlenecks when serving large language models at scale.

"We designed Jalapeño to minimize data movement and communication delays," the company wrote in a blog post presenting the results. "This means that model state, including the KV cache used while generating a response, can be explicitly placed and kept local while the system activates the right combination of compute, memory, and networking for each inference phase."

That last point matters for anyone running production AI workloads. The KV cache — the accumulated context a model holds while generating a response — is one of the biggest memory and bandwidth headaches in modern inference. Keeping it local and explicitly managed, rather than shuffling it across a cluster, is a classic way to claw back latency and power.

Better Than Blackwell — For Now

The headline comparison is against an Nvidia Blackwell system, currently the workhorse of frontier AI inference. Independent analysis published the same day by SemiAnalysis, titled "OpenAI Jalapeño: Better than Nvidia Blackwell," reached a similar conclusion about the chip's early results, and the story quickly became one of the most-discussed items on Hacker News.

But OpenAI executives and analysts alike urge caution. Ho estimated that Jalapeño would deploy at the end of 2026 "in very small volumes," with more significant deployment coming in 2027. By the time Jalapeño reaches full scale, the competition may have advanced significantly — Nvidia's roadmap continues to move, and other hyperscalers including Google, Amazon, and Microsoft are all pushing their own custom silicon.

As TechCrunch noted, benchmarks measured against today's state of the art are a snapshot, not a guarantee. A chip that wins on efficiency in 2026 must keep winning against whatever Nvidia and its rivals ship next year.

Why OpenAI Is Building Its Own Silicon

The strategic logic is straightforward: inference is OpenAI's biggest recurring cost. Every ChatGPT query, every API call, and every agentic task burns compute, and renting that compute from Nvidia at market prices squeezes margins as usage scales. A custom chip tuned to OpenAI's own models — designed jointly with Broadcom and shaped by OpenAI's models themselves — gives the company a way to serve more users per watt and per dollar.

Jalapeño is also planned as a multigenerational platform rather than a one-off part. OpenAI has said it intends to develop AI products, models, chips, and memory together, so that each generation of silicon is co-optimized with the models it will run. If the first-generation results hold up in production, Jalapeño becomes the template for everything that follows.

The stakes extend beyond OpenAI. Nvidia's dominance of AI accelerators has been the single tightest bottleneck in the industry for three years, and every credible alternative — whether from Google's TPU team, Amazon, or now OpenAI — changes the negotiating landscape for everyone building AI infrastructure. Nvidia itself is reportedly warning customers of price hikes of 15 percent or more on AI servers as memory costs surge, which only sharpens the appeal of in-house alternatives.

What Comes Next

For now, Jalapeño remains a pre-production part with promising paper numbers. The real test arrives at the end of 2026, when the first small volumes are deployed, and especially in 2027, when OpenAI expects meaningful scale. Key questions remain unanswered: real-world reliability at datacenter scale, total cost of ownership versus rented Blackwell capacity, and whether the efficiency edge survives against Nvidia's next generation.

Still, the first benchmarks mark a milestone: the first time one of the big AI labs has publicly demonstrated custom inference silicon that appears to beat the incumbent's best on efficiency. For an industry that has treated Nvidia supply as a strategic lifeline, that is a genuine inflection point.

---

Stay Ahead of AI

Get the latest AI news, analysis, and breakthroughs — all in one place.

Read more AI news →