OpenAI has published the first benchmark results for Jalapeño, its custom inference chip, and the numbers are striking: on SemiAnalysis' InferenceX benchmark, the new silicon registered both more tokens per user and more throughput per kilowatt than the currently available state-of-the-art inference processors — a comparison that includes Nvidia's Blackwell systems. The results were presented Tuesday at the Hot Chips conference in Palo Alto, alongside a detailed technical look at how the chip was built. For readers tracking the compute race that underpins every AI model release, it is one of the most consequential hardware datapoints of the year.
"The bottom line is that the results show a very, very significant performance advance over state of the art," said Richard Ho, OpenAI's head of hardware, in a press call reported by TechCrunch. "Jalapeño can serve more AI work per unit of power, while also returning responses more quickly. It's very efficient to serve a lot of customers, but it can also be very low latency."
A Chip Built for the Full Inference Pipeline
Jalapeño was unveiled in June 2026 as OpenAI's first custom silicon, developed in close collaboration with Broadcom, with OpenAI's own AI models assisting in the chip's development process. The partnership itself dates back to October 2025, when OpenAI laid out plans for a multigenerational custom accelerator program alongside the networking giant.
What makes Jalapeño different from a general-purpose GPU, according to OpenAI, is that it was designed in concert with the models it will serve. Because the company controls the full stack — AI products, models, chips, and memory — its engineers were able to attack specific phases of the inference process that often cause friction.
In particular, Jalapeño is designed to minimize delays during the prefill and communication phases of processing, which OpenAI says frequently act as bottlenecks when serving large language models at scale.
"We designed Jalapeño to minimize data movement and communication delays," the company wrote in a blog post presenting the results. "This means that model state, including the KV cache used while generating a response, can be explicitly placed and kept local while the system activates the right combination of compute, memory, and networking for each inference phase."
That last point matters for anyone running production AI workloads. The KV cache — the accumulated context a model holds while generating a response — is one of the biggest memory and bandwidth headaches in modern inference. Keeping it local and explicitly managed, rather than shuffling it across a cluster, is a classic way to claw back latency and power.
Better Than Blackwell — For Now
The headline comparison is against an Nvidia Blackwell system, currently the workhorse of frontier AI inference. Independent analysis published the same day by SemiAnalysis, titled "OpenAI Jalapeño: Better than Nvidia Blackwell," reached a similar conclusion about the chip's early results, and the story quickly became one of the most-discussed items on Hacker News.
But OpenAI executives and analysts alike urge caution. Ho estimated that Jalapeño would deploy at the end of 2026 "in very small volumes," with more significant deployment coming in 2027. By the time Jalapeño reaches full scale, the competition may have advanced significantly — Nvidia's roadmap continues to move, and other hyperscalers including Google, Amazon, and Microsoft are all pushing their own custom silicon.
As TechCrunch noted, benchmarks measured against today's state of the art are a snapshot, not a guarantee. A chip that wins on efficiency in 2026 must keep winning against whatever Nvidia and its rivals ship next year.
Why OpenAI Is Building Its Own Silicon
The strategic logic is straightforward: inference is OpenAI's biggest recurring cost. Every ChatGPT query, every API call, and every agentic task burns compute, and renting that compute from Nvidia at market prices squeezes margins as usage scales. A custom chip tuned to OpenAI's own models — designed jointly with Broadcom and shaped by OpenAI's models themselves — gives the company a way to serve more users per watt and per dollar.
Jalapeño کی منصوبہ بندی بھی یک طرفہ حصے کی بجائے کثیر نسلی پلیٹ فارم کے طور پر کی گئی ہے۔ OpenAI نے کہا ہے کہ وہ AI پروڈکٹس، ماڈلز، چپس اور میموری کو ایک ساتھ تیار کرنے کا ارادہ رکھتا ہے، تاکہ سلیکون کی ہر نسل کو ان ماڈلز کے ساتھ مل کر بہتر بنایا جائے جو یہ چلائے گا۔ اگر پہلی نسل کے نتائج پیداوار میں برقرار رہتے ہیں، Jalapeño اس کے بعد آنے والی ہر چیز کے لیے ٹیمپلیٹ بن جاتا ہے۔
داؤ OpenAI سے آگے بڑھتا ہے۔ AI ایکسلریٹروں پر Nvidia کا غلبہ تین سالوں سے انڈسٹری میں واحد سب سے سخت رکاوٹ رہا ہے، اور ہر قابل اعتبار متبادل - چاہے Google کی TPU ٹیم، Amazon، یا اب OpenAI سے - AI انفراسٹرکچر بنانے والے ہر فرد کے لیے گفت و شنید کے منظر نامے کو بدل دیتا ہے۔ Nvidia خود مبینہ طور پر صارفین کو AI سرورز پر قیمتوں میں 15 فیصد یا اس سے زیادہ اضافے سے خبردار کر رہی ہے کیونکہ میموری کی لاگت میں اضافہ ہوتا ہے، جو صرف اندرون ملک متبادل کی اپیل کو تیز کرتا ہے۔
آگے کیا آتا ہے۔
ابھی کے لیے، Jalapeño کاغذی نمبروں کے ساتھ ایک پری پروڈکشن حصہ بنی ہوئی ہے۔ اصل امتحان 2026 کے آخر میں آتا ہے، جب پہلی چھوٹی جلدیں تعینات کی جاتی ہیں، اور خاص طور پر 2027 میں، جب OpenAI معنی خیز پیمانے کی توقع کرتا ہے۔ اہم سوالات لا جواب ہیں: ڈیٹا سینٹر کے پیمانے پر حقیقی دنیا کی بھروسے، ملکیت کی کل لاگت بمقابلہ کرائے پر لی گئی بلیک ویل صلاحیت، اور کیا کارکردگی کا کنارہ Nvidia کی اگلی نسل کے خلاف زندہ رہتا ہے۔
پھر بھی، پہلے بینچ مارکس ایک سنگِ میل کی نشان دہی کرتے ہیں: پہلی بار بڑی AI لیبز میں سے کسی نے عوامی طور پر کسٹم انفرنس سلکان کا مظاہرہ کیا ہے جو ایسا لگتا ہے کہ آنے والے کی بہترین کارکردگی کو مات دے رہا ہے۔ ایک ایسی صنعت کے لیے جس نے Nvidia کی سپلائی کو اسٹریٹجک لائف لائن کے طور پر سمجھا ہے، یہ ایک حقیقی انفلیکشن پوائنٹ ہے۔
---
AI سے آگے رہیںتازہ ترین AI خبریں، تجزیہ اور کامیابیاں حاصل کریں — سب ایک جگہ پر۔
مزید AI خبریں پڑھیں →