Every answer and its explanation appears here once you have finished the path. Each one then links to the matching glossary entry, where the concept is set out in full with its worked example.
1. A site advertises dual-path power, duplicated cooling and several generators. Which event does this architecture do nothing against, and which underwriting must therefore assess separately?
A loss that compromises the entire building, backup systems included
N+1 or dual-path redundancy protects against the failure of a component inside one building. It does not protect against a peril that makes no distinction between primary and backup systems, since both sit in the same place: fire, flood, or total loss of external power. On 10 March 2021, the fire in building SBG2 at the OVHcloud Strasbourg site made the point: conventional internal redundancy did not prevent the loss of the building, nor the permanent loss of data for clients with no external backup. Two separate resiliences must therefore be read, resilience to routine failures, well captured by the Tier classification, and resilience to catastrophic loss, which depends entirely on a recovery site at a geographic distance.
Glossary entry · redondance-intra-site-limites2. An operator presents a site as "Tier IV" in its submission. Which check changes the risk assessment the most?
Whether the Tier IV covers design, constructed facility or operational sustainability, which are three separate certifications
The Uptime Institute classification runs from Tier I, with no redundancy, to Tier IV, tolerant of any single component failure, with a theoretical annual downtime of a few minutes; Tier III, described as concurrently maintainable, allows maintenance without interrupting service. But the certification separates the design level, the level actually constructed and the level actually operated. An operator can therefore claim a Tier IV design while running a site that was never certified as constructed, and the gap between advertised and real resilience is exactly what underwriting has to measure. The rest of the list matters, but none of those checks moves the business interruption assessment as far.
Glossary entry · classification-tier-uptime3. On 28 February 2017, an outage of Amazon Web Services' S3 storage service interrupted thousands of client sites and applications for roughly four hours. None of those clients had suffered any physical damage. Why does their conventional business interruption policy not respond?
Because conventional business interruption requires covered physical damage to the insured's own property
Conventional business interruption is the consequence of covered physical damage to the insured's own property. Here there is none: the outage sits with a third party, and nothing of the insured's is broken. That is the reason the contingent business interruption clause for a cloud provider exists, since it indemnifies exactly this situation. It shifts the analysis from the insured's site to the provider's, and to the precise technical architecture of the service used, which the end client often does not know. And it creates accumulation of another order: a single event at a dominant provider triggers simultaneous claims across independent insureds in one portfolio, which an interruption confined to a single site never does.
Glossary entry · clause-interruption-fournisseur-cloud4. The engineer behind that outage ran a routine maintenance command with too broad a parameter, removing more servers than intended. How does a property and business interruption policy usually treat that kind of error?
It covers it, ordinary error being part of normal hazard, the difficulty lying in establishing the cause
Ordinary human error is statistically the leading cause of major data center and cloud outages, ahead of hardware failure and ahead of natural catastrophe. It is not the same as sabotage or characterized gross negligence, and it remains covered by property and interruption policies without a specific exclusion. The real difficulty is technical rather than legal: establishing that an outage came from human error rather than from an underlying hardware defect that the error merely exposed, in a field where technical logs may be incomplete or affected by the incident itself. This is why underwriting increasingly examines change management procedures and the existence of test environments isolated from production.
Glossary entry · clause-erreur-humaine-exploitation5. In August 2016, the failure of an electrical control module in a Delta Air Lines data center in Atlanta prevented the transfer to backup generators: thousands of flights cancelled over several days, at an estimated cost of around 150 million dollars. What makes this peril so hard to assess in advance?
The backup chain is almost never exercised under real conditions, so a latent defect can stay invisible for years
The backup chain links two functions: the UPS provides immediate continuity, then the generators take over for the duration. A break at any link, degraded batteries, a faulty control module, a generator that fails to start, a badly synchronized transfer, turns a normally invisible grid outage into a full stop. Yet most tests are run unloaded or only partially, so the defect surfaces only when the chain is genuinely needed. The practical underwriting consequence is clear: the frequency and realism of full-load transfer tests discriminate between sites better than any description of the installation.
Glossary entry · defaillance-onduleur-generateur6. A cyber portfolio holds forty insureds across different sectors and sizes, and underwriting judges it well diversified. All of them host their critical applications in the same data center cluster. What has been mismeasured?
Real diversification, since a cluster turns apparently independent risks into correlated risk
Clustering is rational for each operator taken alone: cheap power, tax treatment, fiber connectivity, proximity to infrastructure already in place. At sector scale it produces an accumulation nobody measures, since each actor sees only its own share. A local event, a regional grid outage or a natural catastrophe, can then hit a significant share of the hosted infrastructure at once, across competing operators with no capital link between them. Loudoun County in northern Virginia, nicknamed Data Center Alley, is the most cited example. Diversification by sector and size therefore says nothing about diversification by technical dependency, and it is the second that decides accumulation.
Glossary entry · concentration-geographique-data-center7. In July 2022, a record heatwave in the United Kingdom exceeded the cooling capacity of Google and Oracle data centers in the London area, causing service outages lasting several hours. What does this peril require underwriting to examine?
The site's thermal design margin against updated climate scenarios rather than historical norms alone
A site is sized for a local temperature range drawn from the history available when it was designed. When heat exceeds that range, the cooling system can no longer remove the heat produced, and the operator must shed load or shut down servers to avoid destroying them: the climate hazard then becomes a direct service interruption with no physical damage. The gap therefore sits between the design basis and the up-to-date climate scenario, and it hits first the sites built to older norms in regions warming faster than anticipated. It is a design parameter, not a policy parameter, which is what makes it invisible to anyone reading only the contract.
Glossary entry · risque-canicule-data-center8. In a data center fire, what decides between a partial loss and a total loss of the site?
Compartmentation between rooms and between buildings
Fire is made worse by the density of energized electrical equipment, by cabling and by flammable material in quantity, and extinguishing is itself a problem: water destroys electronics, and a gas system is not always enough on a fire already developed in a dense room. What bounds the loss is therefore compartmentation: contained to one room it makes a partial loss; spread to the building it makes a total loss. On 10 March 2021, the fire in the SBG2 building at OVHcloud Strasbourg destroyed the building and damaged neighbouring ones, cutting access to millions of hosted sites. Such a loss is both property damage and business interruption, multiplied by the number of hosted clients, and the building's insured value says nothing about that second part.
Glossary entry · incendie-data-center9. Why is a data center dedicated to training artificial intelligence models harder to insure than a traditional center of the same floor area?
Because several determinants degrade at the same time, and no loss history covers these configurations
A center's insurability aggregates several determinants: redundancy level, density per rack, cooling technology and the perils specific to it, grid dependency and autonomy, the site's climate exposure, concentration of value, and replacement lead times for critical equipment. Model training degrades them together: record densities, liquid cooling still poorly documented, sharply rising unit value of graphics processing clusters, supply tensions lengthening restoration times. On top of that sits the absence of loss history on these recent configurations, which makes pricing uncertainty structural rather than cyclical: it does not shrink by spending more time on it, only by accumulating experience.
Glossary entry · assurabilite-data-center