Nvidia and Cerebras are selling performance their customers will (probably) never see
Nvidia and Cerebras are promoting exceptionally high single-user AI inference speeds, but the article argues that most customers are unlikely to experience them in real-world services. Such figures are useful marketing benchmarks, yet commercial operators normally batch many requests together to make inference economical, which reduces the relevance of peak per-user token-generation rates.
Nvidia says its Groq-3-based LPX racks reached 3,400 tokens per second on Gemma 4 31B, around four times Cerebras’ earlier result, while Cerebras claims near-equivalent performance from its new CS-4 systems. Both architectures rely heavily on fast on-chip SRAM, but limited memory restricts batching: the article estimates that a 256-LPU Nvidia rack and a 132GB Cerebras CS-4 rack could each support only about 12 simultaneous 100,000-token sequences for this model before memory is exhausted.
- Peak inference speeds may not translate into practical customer performance.
- Limited memory constrains batch sizes on both systems.
- Commercial AI services prioritise efficient throughput over single-user speed.