Google Cloud outage shows it’s still hard to understand hyperscalers’ real resilience regimes

← Back to the feed

Google Cloud outage shows it’s still hard to understand hyperscalers’ real resilience regimes

The Register · 12 hours ago

A Google Cloud outage last week has highlighted how difficult it remains for customers to understand the true resilience of hyperscale cloud services, according to industry analysts. The incident, caused by an upstream power failure that triggered a cooling failure at a single datacentre in the europe-west4-a zone, disrupted three services for 15 hours, revealing that some cloud offerings depend on a single physical facility even though customers are typically advised to spread workloads across multiple zones for redundancy.

The affected services—VMware Engine, NetApp Volumes and Bare Metal Solutions—were taken offline after an electrical fault on the utility grid disrupted power and cooling equipment, prompting Google to proactively shut down workloads to protect customer data. Google has not explained why the fault caused such disruption or whether on-site backup generation was available, and its full incident analysis remains ongoing. Analysts, including Forrester's Biswajeet Mahapatra, said the real problem is a lack of transparency: customers are rarely told when a managed service relies on a single-datacentre dependency within a zone, leading many to overestimate the redundancy actually built into specialised cloud services across major providers.

  • Google Cloud outage exposed hidden single-datacentre dependency for three services
  • 15-hour disruption in europe-west4-a zone caused by power and cooling failure
  • Analysts say cloud providers lack transparency about true resilience architecture

Business Markets

Read the full article at the source →