🚨 BREAKING:
@Cerebas just stopped pitching wafer-scale as a monolith.
we got to the supernova cerebras event this week and have some thoughts on their CS-4 launch.
it's a beast.
but the CS-4 story is not just about tokens per second ... (which cloaked 120b above 4.4k tps per user!)
the change is the fastest decode silicon in the industry is now built to sit inside somebody else's cluster, with prefill running on whatever is best at prefill.
native disaggregation out of the box - here’s the breakdown...
The CS-4 incorporates:
- 2.4 Tb/s off-wafer per wafer, double CS-3
- RoCE v2 over standard ethernet, no proprietary switch required
- switchless wafer-to-wafer links, 5µs → 2µs
- a claimed path to attention-feedforward disagg on top of prefill-decode
for a machine whose defining constraint is 44GB of SRAM per wafer, I/O is the release valve.
push KV off the wafer and batch goes up without adding capacity. AMD Helios and AWS Trainium are already named as prefill partners.
the compute story is exciting too. Cerebras doubled per-wafer performance on the same 5nm WSE-3 die, no tape-out, by moving power conversion to 0.5mm from the processor. roughly 100x closer than a conventional GPU board. twice the power in, roughly double the clock. doubling frequency on 46,225mm² without a superlinear power blowup is hard engineering.
21.6 to 43.2 PB/s per wafer
3 wafers per rack, up from 2
750 PFLOPS, 129.6 PB/s, 7.2 Tb/s per rack
over 4,400 tok/s/user on gpt-oss-120b
and the rack finally looks like infrastructure.
one power rack, three pluggable backpacks, 50% fewer components, no cable plant, cooling pulled out on the assumption the facility is already liquid. deployment in hours.
what's still open to be solved by next gen:
44GB per wafer did not move and cannot until new silicon. @SemiAnalysis models that 1.6T params at 256k context and 256 concurrent requires roughly 40 WSE systems. basically relegates the system to large buyers. you wouldn’t want to own just one of these.
2x perf at roughly 2x power is flat perf/watt, so the 10x throughput-per-watt claim is quoting a much higher concurrency point than CS-3 was ever sold at. that's a pareto shift, which is exciting, but it's a different claim than the headline in our opinion.
none of that is a knock, the roadmap is telling this story and we are excited about the visibility into next generation launches here at general compute.
the takeaway is the direction they have taken.
We think nobody wins inference by owning the whole rack. it's prefill and decode on the silicon each phase actually needs, assembled by operators who aren't locked to one vendor. the whole decode silicon market is converging towards disaggregated inference.
CS-4 is the biggest vote yet for that architecture, from the company with the most to lose by admitting it.
Final thoughts massive congratulations to the
@cerebras team - more good silicon in the ecosystem is good for everyone building on it.. Can’t wait to see it in production.