Cerebras has taken the wraps off the CS-4, a rack-scale AI inference system built around its new WSE-3 Turbo wafer-scale processor, claiming the machine delivers up to 30 times faster inference than GPU-based systems as the company positions itself as the speed alternative to Nvidia's data center empire.
The launch, announced Tuesday and detailed across the company's product pages and a wave of coverage from Reuters, Bloomberg and The Register, marks Cerebras' most significant hardware refresh since going public — and a pivotal moment for a company whose stock has struggled since its IPO. Investors appear to be paying attention: shares of Cerebras (CBRS) surged on the news, and Investor's Business Daily quoted CEO commentary that the AI inference market is "growing so fast."
For readers tracking the accelerating race to make AI models respond faster, the CS-4 launch is one of the most consequential hardware stories of the quarter, and it lands amid broader shifts in AI industry coverage that continue to reshape the compute landscape.
What's Inside the CS-4
The headline component is the WSE-3 Turbo, an upgraded version of Cerebras' dinner-plate-sized Wafer Scale Engine. According to The Register's technical deep dive, the WSE-3T is not new silicon — it retains the same TSMC 5nm process, 46,225 square millimeter die area, 4 trillion transistors, 900,000 cores and 44 GB of on-wafer SRAM as the two-year-old WSE-3.
Instead, Cerebras achieved its doubling of performance by pushing the existing chip much harder. The WSE-3T doubles sparse FP16 compute to 250 petaFLOPS, doubles dense FP16 performance to an estimated 25 petaFLOPS, doubles memory bandwidth to 43.2 petabytes per second, and doubles I/O bandwidth to 2.4 terabits per second. The Register estimates the company is now clocking the silicon at roughly 2.8 GHz, up from about 1.4 GHz in the previous generation.
The trick, according to Cerebras, is a redesigned power delivery system that sits just 0.5 millimeters from the processor — roughly 100 times closer than the ~50 millimeters typical of conventional GPU boards. That proximity nearly eliminates board-level power loss and allows twice as much power to reach the wafer, enabling the higher operating frequencies behind faster token generation.
The Nexus Rack Architecture
Beyond the chip itself, the CS-4 introduces what Cerebras calls the Nexus Platform Architecture, its first true rack-scale design and a direct answer to Nvidia's NVL72 and AMD's Helios rack systems. Each CS-4 can carry up to three WSE-3T processors housed in self-contained "Wafer-Scale Backpacks" — 3D packages that fold the wafer, power conversion, direct liquid cooling, high-speed I/O and control electronics into a single assembly with 50% fewer components.
The modular design separates the stable power, cooling and network layer from the compute. A Cerebras PowerRack can be installed and facility-qualified before any compute arrives, and the backpacks then slide into place — reducing deployment time from days to hours, according to the company.
Power consumption is substantial. The Register estimates around 46 kW per backpack and a total system draw of 120 to 140 kW — eye-catching numbers that nevertheless look conservative next to the 240 to 250 kW rack systems AMD and Nvidia are expected to ship later this year.
Killing the Switch
Perhaps the most technically interesting choice is the interconnect. Rather than routing wafer-to-wafer traffic through switches — the approach AWS adopted for its Trainium3 accelerators — Cerebras wires its chips into a 2D torus topology where neighboring wafers talk directly to one another. The design cuts wafer-to-wafer latency from five microseconds to just two, and Cerebras says the mesh can support models of up to 50 trillion parameters, comfortably beyond anything that currently exists.
The speed results are striking. On a single CS-4 system, benchmarking firm Artificial Analysis measured up to 4,400 tokens per second per user on gpt-oss-120b — roughly 12 times the ~350 tokens per second of the fastest GPU-based inference services available today. Cerebras also claims the system sustains more than 1,000 tokens per second on models exceeding 10 trillion parameters.
A Disaggregated Partnership Strategy
Notably, Cerebras is no longer trying to run the entire inference pipeline on its own silicon. The company has partnered with AWS and AMD to offload the compute-intensive prompt-processing stage of inference onto Trainium XPUs and Instinct GPUs respectively, leaving its wafer-scale engines to handle the latency-sensitive decode phase where their massive SRAM reserves shine.
The CS-4 launch also extends Cerebras' collaboration with OpenAI. Yahoo Finance reported the unveiling alongside new OpenAI and AMD partnerships aimed at accelerating AI inference — building on the Ultrafast mode the companies introduced earlier in August, which brought GPT-5.6 Sol to users at up to 750 tokens per second.
Caveats Behind the Headline Numbers
Industry analysts urge some caution on the marketing figures. The Register notes that Cerebras' headline 250 petaFLOPS relies on sparsity, which generally does not benefit large language model inference, and that peak memory bandwidth figures are largely theoretical since the WSE-3 lacked the compute to saturate its own SRAM. The publication also questioned why Cerebras doubled performance rather than increasing SRAM capacity, which has not meaningfully grown since the WSE-2 launched five years ago — a curious choice in an era of disaggregated inference where decode accelerators benefit most from memory.
Still, the market opportunity is real. As AI chatbots and agents demand ever-faster response times, inference speed has become a genuine competitive differentiator, and Cerebras has now demonstrated it can iterate its unconventional wafer-scale technology on a commercial cadence.
The first CS-4 shipments begin this quarter, with the first systems expected online before the end of September. Whether that's enough to dent Nvidia's dominance remains to be seen — but for the first time in a while, the fastest AI inference in the industry has a new address.
Stay Ahead of AI
Hardware launches like the CS-4 move markets and redefine what AI systems can do. Read more AI news to keep up with every development.
Explore the latest AI coverage →