Nvidia and Cerebras are selling performance their customers will (probably) never see

← Back to the feed

Nvidia and Cerebras are selling performance their customers will (probably) never see

The Register · 6 hours ago

Nvidia and Cerebras are promoting exceptionally high single-user AI inference speeds, but the article argues that most customers are unlikely to experience them in real-world services. Such figures are useful marketing benchmarks, yet commercial operators normally batch many requests together to make inference economical, which reduces the relevance of peak per-user token-generation rates.

Nvidia says its Groq-3-based LPX racks reached 3,400 tokens per second on Gemma 4 31B, around four times Cerebras’ earlier result, while Cerebras claims near-equivalent performance from its new CS-4 systems. Both architectures rely heavily on fast on-chip SRAM, but limited memory restricts batching: the article estimates that a 256-LPU Nvidia rack and a 132GB Cerebras CS-4 rack could each support only about 12 simultaneous 100,000-token sequences for this model before memory is exhausted.

  • Peak inference speeds may not translate into practical customer performance.
  • Limited memory constrains batch sizes on both systems.
  • Commercial AI services prioritise efficient throughput over single-user speed.

Software

Read the full article at the source →