AI training pushes data centre networks to their limits
AI training is increasingly spread across data centres, as operators seek the power, space and planning flexibility that a single site may lack. This matters because training large models depends on thousands of accelerators exchanging information in sync, making the network between sites part of the computing system.
The article says Google trained Gemini across locations, Microsoft linked facilities in Wisconsin and Georgia, and AWS connected clusters for Anthropic; Meta and other providers are also building cross-site links. Cisco estimates current training can require tens of thousands of GPUs, while Epoch AI researchers project the largest frontier runs could need 4–16GW by 2030. During training, accelerators exchange numerical data such as gradients, and congestion or delays can leave others waiting or, in severe cases, force a return to a saved checkpoint.
- AI training is spreading across geographically separate data centres.
- Synchronous GPU work makes network delays a computing bottleneck.
- Frontier training runs could require 4–16GW by 2030.
New here? Start with this
Training artificial intelligence models requires enormous computational power. Large models need thousands of specialised processors working together, demanding far more electricity, cooling and physical space than any single data centre can provide. Technology firms including Google, Microsoft, Amazon and Meta are spreading this training across multiple facilities in different locations to access the resources they need.
When these processors are split across different data centres, the network connecting them becomes part of the computing system itself. The processors must constantly exchange calculations in perfect synchronisation—delays or congestion can force the system to pause and restart, wasting time and energy. The connections between sites are therefore as critical to success as the hardware.
The scale is growing rapidly. Cisco estimates current systems require tens of thousands of processors, whilst research groups project that the largest training runs by 2030 could demand 4 to 16 gigawatts of electricity. As AI training becomes more ambitious, the pressure on networks linking distant data centres will intensify.
Read the full article at the source →
Originally published by The Register as “Growing pains: how distributed AI training changes the network between datacenters”.