Cerebras CS-4: 750 PFLOPs Across Three WSE-3 Turbo Wafers
Cerebras CS-4wafer-scaleинференс ИИ
What Cerebras built into CS-4
For me, the most important part of CS-4 is not the PFLOPs figure, but the attempt to eliminate the movement of weights between external memory and compute units. CS-4 is a rack-scale system built around three WSE-3 Turbo wafers, not simply another standalone accelerator. On the CS-4 product page and in its launch announcement, Cerebras cites 750 PFLOPs of AI compute.
Each wafer includes 44 GB of SRAM, where the architecture keeps model weights. Claimed memory bandwidth is 129.6 PB/s, while the on-chip fabric is rated at 160.5 PB/s. System I/O is listed at 7.2 Tb/s with 2 microseconds of latency.
This is where the story gets interesting: Cerebras claims up to a 30-fold inference speedup over GPU-based systems. In a direct comparison using GPT-OSS-120B and identical prompts, the company reports more than 4,400 tokens per second per user. These are vendor figures from launch materials, not a universal guarantee for every model or serving configuration.
The hardware price has not been published. As of August 2026, Cerebras's public pricing page covers hosted services and APIs rather than CS-4 itself. That means an honest comparison of total cost of ownership and performance against a GPU cluster is not yet possible from public data.
Why the architecture matters more than the headline number
If the claimed results hold up on real workloads, CS-4 primarily changes the inference architecture for large models. Rather than constantly moving weights between accelerators, the system relies on keeping them directly in wafer SRAM. The benefit should be most visible when memory latency matters more than nominal compute throughput.
I would first test not peak speed but result stability across different prompt lengths, batching patterns, and numbers of concurrent users. Model support, software integration, and real operating costs are also open questions. A claim of up to 30 times faster inference sounds impressive, but without identical methodology and pricing it remains only part of the picture.
CS-4's strongest feature is also its main risk: the specialized wafer-scale architecture must prove its advantage not in one showcase run, but in day-to-day model serving. That is where it will become clear whether this is a new infrastructure class or a very fast but narrow tool.