Cerebras and OpenAI have given the world an early look at Ultrafast Mode, a new service tier that launches first in the OpenAI API and is powered by Cerebras hardware. According to the Cerebras announcement published August 13, 2026, GPT-5.6 Sol running on Ultrafast delivers up to 750 output tokens per second — with no compromise in quality.

The tier is initially available to a select group of customers, with access expanding over time. The announcement, written by Cerebras' Joyce Er, positions the offering as a resolution to a trade-off that has defined AI product development for years: builders had to choose between speed and intelligence, because larger and smarter models incur higher computational and data-movement costs that slow down response times. Readers following the race to accelerate frontier models can track the latest developments on AI Buzz Wire.

How Fast Is Ultrafast, Exactly?

The raw throughput number is striking on its own, but Cerebras also supplied comparative context. According to output speeds reported by Artificial Analysis, an independent benchmarking service, GPT-5.6 Sol on Ultrafast mode runs:

  • 11x faster than Fable 5
  • 5x faster than Opus 4.8 on Fast mode

To stress-test the claim, Cerebras ran the model head-to-head against popular competitors on Humanity's Last Exam (HLE), a challenging benchmark consisting of 2,500 questions that are typically answerable only by people holding PhDs in fields such as chemistry, economics, and literature.

The result: GPT-5.6 Sol on Ultrafast answered all 2,500 HLE questions in 11 hours and 11 minutes. Claude Fable 5 needed 78 hours and 27 minutes — more than three days of continuous compute — to arrive at the same conclusions. In other words, Ultrafast worked through the frontier of human knowledge in a single working day, achieving comparable accuracy nearly 7x faster.

Cerebras noted its benchmarking methodology: GPT-5.6 Sol Ultrafast was tested with Codex on xhigh reasoning on July 10, while Claude Fable 5 was tested with Claude Code on xhigh reasoning July 13–15.

Real-World Workloads, Not Just Benchmarks

Speed benchmarks can be misleading if quality degrades along the way, which is why Cerebras also highlighted results on GDP-Val, a benchmark for economically valuable knowledge work tasks. On GDP-Val, Ultrafast delivered a 5.6x end-to-end speedup with no quality degradation, according to tests Cerebras ran on July 31, 2026 using GPT-5.6 Sol and GPT-5.6 Sol Ultrafast on medium reasoning within Codex.

The company says GPT-5.6 Sol is OpenAI's best model yet for legal briefs, financial models, and engineering reports — domains where output quality and turnaround time both carry direct financial consequences.

Why Speed Changes the Agent Equation

The more consequential shift may be in how AI agents operate. Faster inference changes what is possible for individuals and organizations, because agents can be placed on the critical path of problems where every second counts.

"With GPT-5.6 Sol Ultrafast, Cerebras enables AI that keeps up with how you think, code, and collaborate," said Rohan Varma, Product at OpenAI, in a statement quoted in the announcement. "We're excited to see how workflows and applications are transformed by Ultrafast inference."

Cerebras outlined several high-stakes scenarios where the speedup matters:

Incident response

Companies operating web services can use Ultrafast to root-cause and address production outages, preserving customer trust, preventing lost revenue, and saving downtime minutes against their SLAs.

Cybersecurity

In adversarial, high-stakes cyberattacks, security teams must quickly detect and respond to bad actors to contain catastrophic losses — a task where inference latency directly affects damage control.

Agent productivity

Ultrafast enables new modes of working with agents by delivering real-time insights and updates, reducing the need to context-switch across parallel sessions.

"Whereas formerly I might have to wait a couple minutes for a task to finish, it now finishes for me before I even have the opportunity to context-switch. It makes me way more productive," said Jeffrey Wang, a researcher at OpenAI, in the announcement.

The Strategy Behind the Partnership

For Cerebras, the deal represents another milestone in its effort to commercialize wafer-scale acceleration for AI inference. For OpenAI, Ultrafast Mode adds a premium speed tier to its API at a moment when inference latency has become a competitive differentiator — particularly for agentic applications that chain many model calls together, where per-call delays compound into minutes of waiting.

The companies frame the tier as a persistent edge for organizations using frontier AI to respond quickly to incoming information, pairing Ultrafast processing for time-sensitive work with standard processing for commodity tasks that can be parallelized in the background.

Access is rolling out gradually. Interested developers can monitor the OpenAI API documentation for availability, as Cerebras says access will expand over time from the initial select group of customers.

Stay Ahead of AI

For more coverage of the hardware and infrastructure race reshaping artificial intelligence, bookmark AI Buzz Wire and follow our AI hardware news feed.

Read more AI news →