Cloud reliability data & benchmarks

47

major public incidents in the corpus, 2017–2025

3h 24m

median major-incident duration

11h

90th-percentile duration - plan for the tail

Major incidents by provider

Raw counts favor no one: AWS leads partly because it discloses more and runs more surface area. Frequency differences are smaller than duration and blast-radius differences.

AWS19 major incidents
Microsoft Azure16 major incidents
Google Cloud12 major incidents

Median incident duration by provider

Median recovery ranges from under 3 hours (Google Cloud) to nearly 4 (Azure). The medians hide the tail: the corpus's longest events all exceeded 10 hours.

AWS3h 37m median
Microsoft Azure3h 46m median
Google Cloud2h 48m median

Root causes of major incidents

Configuration and change error dominates at 45% - the industry's biggest reliability lever is change safety, not more redundancy. External attack accounts for just 6%.

Configuration / change error45%
Software defect21%
Capacity / overload13%
Network / hardware failure11%
External attack (DDoS)6%
Other / undisclosed4%

Flagship-region and global-layer concentration

Failures concentrate where control planes live. us-east-1 features in 63% of AWS majors; Azure and GCP concentrate instead in global layers (identity, WAN, API management) that cross every region simultaneously.

AWS incidents involving us-east-163%
Azure incidents involving global services44%
GCP incidents involving global layers50%

Major incidents per year (all providers)

Five to seven majors a year, remarkably stable - reliability investment and complexity growth appear to be cancelling out.

2019incidents/year2025

Want the incident-level tables? They ship with the full 2026 report. For how incidents enter the corpus and the known limitations, read the methodology.