Cerebras Systems has introduced the CS-4, a rack-scale computing platform built around its newly launched Wafer Scale Engine 3 Turbo (WSE-3T) chip. The company said the CS-4 delivers 750 petaflops of AI compute, positioning it as a significant leap forward in high-performance AI infrastructure.
The CS-4 is the first product in Cerebras' Nexus rack-scale platform architecture. According to the company, it is twice as fast as the previous-generation CS-3 and offers a 30x tokens-per-second-per-user advantage over GPU-based systems. The system features a modular design that separates compute, power, and I/O, with a "rear-mounted backpack" that attaches vertically to the power array, integrating power conversion, direct liquid cooling, and high-speed I/O in a single package built around the wafer.
By decoupling compute from power supplies, Cerebras said the manufacturing process is simplified and deployment time is reduced from days to hours. The CS-4 provides 7.2 terabits per second of I/O bandwidth and 129.6 petabytes per second of memory bandwidth, and is designed to support very large clusters and models exceeding 50 trillion parameters.
The WSE-3T chip at the heart of the CS-4 contains 4 trillion transistors and 900,000 AI-optimized cores, spread across 46,225 square millimeters of silicon. It delivers 250 petaflops of AI compute per wafer, 43.2 petabytes per second of memory bandwidth, and 44GB of SRAM. Cerebras said the WSE-3T, like its predecessor the WSE-3, is the "largest AI processor ever built."
"In AI, speed is productivity," said Andrew Feldman, CEO and co-founder of Cerebras. "Historically, fast inference meant using smaller and less capable models. Cerebras CS-4 delivers industry-leading speeds on the largest frontier models, fundamentally changing the paradigm. Every aspect of the design has been optimized to deliver the highest speeds with massive throughput. With the CS-4, AI is so fast that it fundamentally reshapes product experiences."
The announcement underscores the growing competition in the AI hardware sector, where custom silicon and specialized architectures are increasingly challenging traditional GPU dominance. Cerebras' focus on inference speed and scalable cluster deployment suggests a broader industry shift toward systems optimized for both training and real-time AI workloads.
First shipments of the CS-4 are expected to begin in the current quarter, the company said.