OpenAI has unveiled its first custom-built inference processor, a chip named Jalapeño that was designed and manufactured in collaboration with Broadcom. The announcement marks the ChatGPT maker's most significant step yet toward owning the full hardware stack behind its AI services, and it intensifies the competitive pressure on Nvidia, whose GPUs currently power most large language model deployments.

Revealed on June 24, 2026, the processor was designed specifically for the demands of OpenAI's inference systems — the process of running pre-trained AI models in response to user requests. According to TechCrunch, which reported the launch, the company said that its own AI models even assisted in the chip's development. While Jalapeño is still being tested, OpenAI claims early results show significantly better performance-per-watt than current state-of-the-art alternatives. For more context on this story, see our ongoing more AI stories.

A Chip Built for Inference, Not Training

The distinction between training and inference is central to understanding why OpenAI built Jalapeño. Training a frontier model — the months-long, compute-intensive process of teaching a neural network from scratch — will continue to rely heavily on Nvidia hardware. Inference, by contrast, happens every time a user sends a prompt, and it accounts for the vast majority of the ongoing cost of running a service like ChatGPT.

In its announcement, OpenAI emphasized the chip's low operating cost when running real-time coding models, the kind of agentic workloads it has been pushing through products like Codex. TechCrunch reported that even modest reductions in inference costs could meaningfully improve the company's economics, given the enormous volume of requests its platforms handle daily.

The design partnership with Broadcom was officially announced in October 2025, though OpenAI's silicon ambitions had been rumored for far longer. Reuters, Bloomberg, CNBC, and the Wall Street Journal all independently confirmed the June 24 unveiling, signaling the breadth of industry attention the launch commanded.

The Full-Stack Bet

OpenAI framed Jalapeño as part of a broader strategy to control every layer of its infrastructure. The company wrote that it is "not only developing frontier models or building products on top of them; it is designing the infrastructure underneath them: chip architecture, kernels, memory systems, networking, scheduling, deployment systems, and product experience."

OpenAI president Greg Brockman explained the company's rationale on an in-house podcast shortly after the Broadcom partnership was first disclosed. "We have a deep understanding of the workload," Brockman said. "We've really been looking for specific workloads that are underserved, [and asking] how can we build something that will be able to accelerate what's possible?"

Because OpenAI operates across the entire stack, it argues, each layer can be optimized toward a single goal: making its models faster, more reliable, and more affordable for users.

Joining Google and Amazon in Custom Silicon

OpenAI is hardly the first AI company to build its own accelerators. Google has long designed its Tensor Processing Units (TPUs) to serve its models, while Amazon has developed its Trainium and Inferentia chips for AWS customers. These so-called AI accelerators are silicon tuned specifically to speed up machine learning workloads, and they have become a competitive necessity for any hyperscaler operating at OpenAI's scale.

The motivation is straightforward: general-purpose GPUs from Nvidia are powerful but expensive, and they tie operators to a single supplier whose pricing and availability can constrain growth. A custom inference chip gives OpenAI more control over cost, supply, and performance for the workloads that matter most to its bottom line.

What Remains on Nvidia

It would be a mistake to read Jalapeño as a complete break from Nvidia. The most performance-intensive tasks in the AI pipeline — particularly the pre-training of frontier models — still demand the kind of raw, general-purpose throughput that Nvidia's GPUs excel at, and OpenAI remains a major Nvidia customer. Jalapeño instead targets the narrower but enormous workload of inference, where a purpose-built chip can deliver outsized savings.

That narrow focus is also what makes the chip credible. Rather than attempting to replace Nvidia across the board, OpenAI is competing only where it believes its deep knowledge of its own workloads gives it an edge — the same logic that drove Google and Amazon to build specialized accelerators in the first place.

A Shift in the Semiconductor Landscape

The launch arrives amid a broader reshuffling of the AI semiconductor landscape. Memory-chip giant SK Hynix has filed for an estimated $29 billion US listing to capitalize on surging demand for AI memory, while companies from Amazon to Qualcomm have been racing to build or acquire custom silicon. OpenAI's entry into chip design adds one more well-funded competitor to a market that Nvidia has dominated.

For now, Jalapeño remains in testing, and OpenAI has not disclosed detailed performance benchmarks or a deployment timeline. But the signal is clear: the company that built ChatGPT now intends to build the hardware that runs it, and it is willing to spend billions to do so.

Whether Jalapeño delivers on its early performance-per-watt promises will determine how much leverage OpenAI gains over its costs — and how much of the AI infrastructure market eventually shifts away from a single dominant chipmaker.

---

Stay Ahead of AI

Get the latest AI news, analysis, and breakthroughs — all in one place.

Read more AI news →