Executive Overview

Cerebras Systems has unveiled the third iteration of its wafer-scale engine, the CS-3. By manufacturing an entire 300mm silicon wafer as a single contiguous processor, Cerebras circumvents the physical interconnect bottlenecks that plague traditional multi-GPU clusters.

Silicon Specifications & Memory Bandwidth

The CS-3 packs 4 million cores and 44 gigabytes of on-chip SRAM, delivering an unprecedented 21 petabytes per second of memory bandwidth. Because all model weights and intermediate activations reside directly on-chip, communication latency between cores is measured in nanoseconds.

Key Architectural Metrics

  • Peak AI Compute: 125 Petaflops of FP16 sparse performance.
  • Cluster Scalability: Up to 2,048 CS-3 systems can be networked to train 24-trillion-parameter foundation models without code partitioning.
  • Power Efficiency: Eliminates external DDR/HBM memory controller power overheads.