Outage caused by a configuration error in a major operator's own internal network, able to cut access to its services with no hardware failure at all.
Internal network failure at a hyperscaler refers to an outage caused not by physical damage but by a configuration error in the operator's internal logical network, in particular routing protocols such as BGP, which govern how the different parts of a global infrastructure communicate with each other. A poorly validated configuration change can cut access to an operator's services worldwide, even as every one of its data centers remains physically intact and powered. This peril illustrates a feature specific to very large digital infrastructures: their scale and automation, which normally make them resilient to local hardware failures, in turn expose them to software failures capable of propagating across the entire system within minutes, with no guarantee that immediate manual intervention can stop it. Recovery can also be slowed if the remote management tools themselves depend on the network that has gone down, forcing engineers to intervene physically on site. For insurers, this peril reinforces the accumulation risk already identified for hyperscalers, demonstrating that a purely software event can have a reach as broad as a physical disaster.
On October 4, 2021, a configuration error in Facebook's internal routing network cut off global access to Facebook, Instagram and WhatsApp for about six hours, and even prevented some employees from physically entering data centers whose badge systems depended on the same network.
internal network outage, panne BGP, défaillance configuration réseau, hyperscaler network failure