

Introducing CS-4: The Fastest AI Accelerator in the Industry Learn more >>
The Fastest AI Just Got Faster. Introducing the all new Cerebras CS-4, a revolutionary rack-scale solution that delivers up to 30x faster inference compared to GPUs, enhanced economics, and a simple path to deploy hyperscale capacity. It is the architecture for frontier AI.
Each wafer delivers up to 2x the speed of the previous generation
Scales massive models and enables heterogeneous, disaggregated inference
Enables rapid deployment in hyperscale datacenters
Powered by WSE-3 Turbo, CS-4 delivers up to 30x faster inference compared to GPU systems, setting a new record for the fastest inference available in production.
The CS-4 solution shifts the inference Pareto frontier, delivering up to 10x more throughput per watt than CS-3 while generating tokens up to 30x faster than production GPU systems. The result is a system designed to deliver both throughput and interactivity.
By reducing wafer-to-wafer interconnect latency to 2 microseconds, CS-4 delivers more than 1,000 tokens per second on models exceeding 10 trillion parameters, preserving interactive decode performance at unprecedented scale.
CS-4 is the first iteration of the new Cerebras Nexus Platform Architecture. It is built around a modular concept with three foundational elements: Compute, Power, and I/O – each with significant innovation to simplify manufacturing, deployment, maintenance, and upgrades.
Cerebras has fundamentally re-imagined the server. Each Wafer-Scale Backpack is a self-contained assembly that folds the wafer, power conversion, direct liquid cooling, high-speed I/O, and control electronics into a compact 3D package with 50% fewer components. This design simplifies manufacturing and reduces deployment time from days to hours.
With power delivery just 0.5 millimeters away from the processor - roughly 100x closer than the roughly 50mm of conventional GPU boards - CS-4 nearly eliminates board-level power loss. This enables the delivery of twice as much power to the WSE-3T, enabling higher operating frequencies and faster token generation.
CS-4 introduces a new programmable I/O subsystem that doubles I/O bandwidth and reduces latency, benefitting both aggregated and disaggregated solutions. The Wafer I/O Module also enables wafers to be linked within and across racks without a switch, for wafer-to-wafer latency as low as two microseconds that is key to interactivity for models with tens of trillions of parameters.
CS-4 runs on WSE-3 Turbo—the world’s largest and fastest AI processor. Its four trillion transistors and 900,000 AI cores deliver 250 PFLOPS of compute and 43.2 petabytes per second of memory bandwidth.
Twice the compute. Twice the bandwidth. Less than half the latency. A massive leap in AI speed and throughput.
How is CS4 different from CS-3? How fast is CS-4? How much throughput does CS-4 provide? Why is CS-4 well suited for agentic AI? What is a Wafer-Scale Backpack? How does the Nexus Platform Architecture simplify hyperscale deployment? What models and inference architectures does CS-4 support? Follow
Performance comparisons are based on third-party benchmarking or internal testing. Observed inference speed improvements versus GPU-based systems may vary depending on workload, configuration, date and models being tested.
Hacker News
news.ycombinator.com