Cerebras Systems and OpenAI have shared an early look at Ultrafast Mode, a new service tier launching first in the OpenAI API and powered by Cerebras wafer-scale technology. The partnership enables GPT-5.6 Sol to run at up to 750 output tokens per second, delivering frontier-level intelligence at real-time speeds. For the latest AI hardware news, visit AI Buzz Wire.

Eliminating the Speed-Intelligence Tradeoff

AI builders have long faced a fundamental tradeoff: as models scale up in size and intelligence, they incur higher computational and data movement costs, slowing down response times. Developers were forced to choose between waiting for high-quality results or accepting inferior results in less time.

GPT-5.6 Sol on Ultrafast Mode is designed to resolve this dilemma. According to Cerebras, the service runs 11 times faster than Anthropic's Claude Fable 5 and 5 times faster than Opus 4.8 on Fast mode, based on output speeds reported by Artificial Analysis.

Humanity's Last Exam: The Speed Test

Cerebras put Ultrafast to the test against competing models on Humanity's Last Exam (HLE), a challenging benchmark consisting of 2,500 questions typically answerable only by those holding PhDs in fields such as chemistry, economics, and literature.

In evaluations conducted by Cerebras, GPT-5.6 Sol on Ultrafast Mode answered all 2,500 HLE questions in 11 hours and 11 minutes. Claude Fable 5 required 78 hours and 27 minutes, more than three days of continuous compute, to arrive at the same conclusions. In other words, Ultrafast processed the frontier of human knowledge in a single working day, achieving comparable accuracy nearly seven times faster.

Benchmarking was performed by Cerebras using GPT-5.6 Sol Ultrafast with Codex on high reasoning settings on July 10, and Claude Fable 5 with Claude Code on high reasoning settings from July 13 to 15.

Real-World Business Impact

On GDP-Val, a benchmark for economically valuable knowledge work tasks, Ultrafast delivered a 5.6x end-to-end speedup with no quality degradation. This demonstrates how faster inference can accelerate economically valuable work without sacrificing output quality.

Cerebras envisions Ultrafast as a tool for scenarios where every second matters. Companies operating web services could use the technology to root-cause and address production outages more quickly, preserving customer trust and preventing lost revenue. In high-stakes cybersecurity situations, security teams could leverage Ultrafast to detect and respond to attacks rapidly.

The Technology Behind the Speed

The performance advantage stems from Cerebras' Wafer-Scale Engine architecture, which is purpose-built for frontier AI workloads. The core challenge of fast frontier inference is a data movement problem: on traditional GPUs, inference on large models is bottlenecked by memory bandwidth, as model weights must be repeatedly transferred between on-chip memory and off-chip storage to generate successive tokens.

Cerebras takes a fundamentally different approach. Each wafer-sized chip packs 44 GB of SRAM directly on the silicon. Model weights stay on-chip, and tokens flow uninterrupted through model layers pipelined across wafers. This eliminates the inefficient data movement that slows GPU-based inference and scales smoothly with model size, paving the way for a continued speed advantage on future frontier models.

Availability and Market Positioning

GPT-5.6 Sol on Ultrafast Mode is currently available in a limited preview to a select group of customers, with access expanding over time as capacity grows. Cerebras describes the partnership as powering the next wave of AI innovation and raising the ceiling for what organizations can accomplish with responsive AI.

The collaboration represents a significant milestone for Cerebras, which has long pursued wafer-scale computing as an alternative to the GPU-dominated AI hardware market. By partnering directly with OpenAI to accelerate one of the most capable frontier models available, Cerebras is making a strong case that its architecture offers a genuine competitive advantage for high-speed inference workloads.

As models continue to grow larger and more intelligent, the ability to serve them at real-time speeds could become a defining competitive factor in the AI infrastructure market. The Cerebras-OpenAI partnership signals that the race for inference speed is now as important as the race for model intelligence.

Stay Ahead of AI

For more AI hardware news, model performance analysis, and industry developments, bookmark AI Buzz Wire.

Read more AI news →