← Back to the feed

AI training now spans data centres as power becomes scarce

The Register ·

Training the largest AI models now requires infrastructure to span multiple data centres rather than operating from a single site, as power availability has become a critical physical constraint. This shift creates new and demanding networking challenges that traditional datacenter interconnect systems were not designed to handle.

Distributing training workloads across geographically separate facilities demands synchronous data flows with minimal packet loss and tightly coordinated GPU communication, where bottlenecks in one location can cascade across the entire job. Cisco's approach, presented by Rakesh Chopra, emphasises the "Scale-Across" imperative, integrating high-speed coherent optics, advanced buffering, programmable silicon, and security protocols such as MACsec and IPsec to make distant facilities behave as a single deterministic computing system whilst managing power efficiency between networking and GPU resources.

  • AI model training now spans multiple data centres due to power constraints
  • Networks must handle synchronous GPU flows with minimal packet loss
  • Distributed infrastructure requires treating separate sites as one computing system

New here? Start with this

Training very large AI models now requires so much power that no single data centre has enough to handle the job. Companies are spreading the training work across multiple facilities in different locations, which allows them to access more power and avoid overloading any one site.

When a training job runs across distant data centres, those facilities must communicate constantly with almost no delays or errors, as if they were operating as a single computer. Any slowdown or problem at one location immediately affects the entire project, making reliable high-speed connections between data centres essential but technically very demanding.

Technology companies are developing new networking systems to link distant data centres together, using faster connections, smarter data management, and security measures to make them function as a unified system. These solutions aim to keep the training process running smoothly whilst also managing power consumption efficiently across both the network equipment and the computing resources.

Both sides, in good faith

The strongest fair case each way — we don't pick a winner.

The case for

The shift to distributed training across multiple data centres represents necessary infrastructure evolution to sustain progress in artificial intelligence. As power availability at single locations becomes the limiting factor, distributing workloads across geographically separate facilities is a pragmatic engineering solution that enables continued research and capability development. The sophisticated networking solutions being deployed represent genuine human ingenuity in solving complex coordination problems, and this technological advancement benefits society by keeping pace with computational demands.

The case against

This development exemplifies the escalating resource costs of ever-larger AI models that few genuinely need, prioritising capability expansion over sustainability and efficiency. The requirement for high-speed inter-datacenter networking adds substantial power consumption and infrastructure complexity that could be avoided by focusing research on smaller, more efficient models or reconsidering whether such massive training runs are justified in terms of societal benefit. Rather than enabling this fundamentally unsustainable trajectory by deploying more sophisticated infrastructure, society should question whether unlimited scaling reflects responsible stewardship of energy and natural resources.

AI Cricket Sport Technology

Read the full article at the source →

Originally published by The Register as “When one datacenter is no longer enough”.