47
major public incidents in the corpus, 2017–2025
3h 24m
median major-incident duration
11h
90th-percentile duration - plan for the tail
Raw counts favor no one: AWS leads partly because it discloses more and runs more surface area. Frequency differences are smaller than duration and blast-radius differences.
Median recovery ranges from under 3 hours (Google Cloud) to nearly 4 (Azure). The medians hide the tail: the corpus's longest events all exceeded 10 hours.
Configuration and change error dominates at 45% - the industry's biggest reliability lever is change safety, not more redundancy. External attack accounts for just 6%.
Failures concentrate where control planes live. us-east-1 features in 63% of AWS majors; Azure and GCP concentrate instead in global layers (identity, WAN, API management) that cross every region simultaneously.
Five to seven majors a year, remarkably stable - reliability investment and complexity growth appear to be cancelling out.
Want the incident-level tables? They ship with the full 2026 report. For how incidents enter the corpus and the known limitations, read the methodology.