Cerebras CS-4 rack systems juice their dinner-plate-sized AI chips for every last drop of AI perf
Cerebras has unveiled its next-generation Wafer Scale Engine chip, the WSE-3T, along with new Nexus rack systems, aiming to boost AI inference throughput per watt by up to tenfold compared with its previous generation. Rather than introducing new silicon, the "Turbo" chip uses the same process technology, wafer size and transistor count as its predecessor but pushes far more power through it, reportedly nearly doubling clock speed, which the company says enables faster token generation for AI workloads.
The WSE-3T offers 250 petaFLOPS of sparse AI compute, 44GB of on-chip SRAM, 43.2 petabytes per second of memory bandwidth and 2.4 Tbps of connectivity, roughly double the specifications of the WSE-3. However, much of this headline performance relies on sparsity, which typically does not benefit large language model inference, so real-world dense FP16 performance is likely closer to 25 petaFLOPS. Notably, Cerebras has also partnered with AWS and AMD to offload compute-heavy prompt processing onto their Trainium and Instinct chips respectively, with its own wafer-scale accelerators now focused primarily on the memory-intensive decode stage of inference, allowing far fewer chips to serve very large models.
- Cerebras unveils WSE-3T chip and Nexus racks, doubling per-chip AI performance.
- New chip is same silicon as before, just run at much higher power and clock speed.
- Cerebras now partners with AWS and AMD, focusing its chips on decode-stage inference.