Cloud Reliability Research & Data
The open benchmark dataset behind our numbers - AWS, Azure and Google Cloud incident frequency, duration, root cause and regional risk - plus platform-by-platform reliability insights. The written analysis lives in the 2026 report.
- Corpus
- 47 major incidents
- Coverage
- AWS · Azure · GCP
- Period
- 2017–2025
- Last updated
- June 30, 2026
01 · Benchmark data
The open corpus behind our figures: 47 major public incidents across AWS, Azure and Google Cloud, 2017–2025, normalised per our severity methodology. Every chart sits next to a text summary so the numbers stay quotable.
47
major incidents analysed, 2017–2025
3.4h
median duration of a major incident
11h
90th-percentile duration
02 · Reliability positioning
Each provider plotted on demonstrated reliability against operational maturity; bubble size reflects market presence. All three qualify as Leaders or Strong Performers - the gaps are in execution, not capability.
- GCP - LeadersReliability 89 · Maturity 82 · Presence 62
- AWS - LeadersReliability 83 · Maturity 88 · Presence 95
- Azure - Strong PerformersReliability 66 · Maturity 71 · Presence 84
03 · Where downtime comes from
Across all three providers, networking and DNS remained the single largest root-cause category - a pattern that has held for three years running. Software deploys are the fastest-shrinking category as progressive rollout and automated rollback mature.
- Networking / DNS31%
- Power & Hardware22%
- Software Deploys19%
- Capacity / Scaling16%
- Configuration12%
04 · Research insights
Reliability profiles of the platforms teams depend on most - each grounded in public incident records and reporting from CRN, TechTarget, the Futurum Group and Time.
Cloudflare Reliability: The Edge Dependency Behind Half the Web
Cloudflare sits in front of a huge share of the internet, so its rare control-plane and config outages have an outsized blast radius. A look at the pattern and how to design for it.
Jul 16, 2026 · 9 min readDatadog Reliability: Why Your Monitoring Must Not Share Fate With Your App
Observability platforms fail too, as the March 2023 multi-region Datadog outage showed. When your monitoring goes dark at the same time as your app, you are flying blind.
Jul 14, 2026 · 9 min readStripe Reliability: Designing Checkout for Graceful Degradation
Stripe is the payment backbone for a large share of online commerce, so a payments incident can stop revenue cold. How to keep checkout resilient when the payment layer degrades.
Jul 11, 2026 · 9 min readSlack Reliability: What a Decade of Outages Teaches About Real-Time Uptime
Slack’s incident history - from the 2021 new-year outage to routine degradations - and what it reveals about running always-on collaboration at scale.
Jul 8, 2026 · 9 min readOpenAI API Reliability: Uptime, Capacity, and the Incident Patterns Behind ChatGPT
How OpenAI’s API and ChatGPT have held up under explosive demand - the recurring capacity and degradation patterns, and what they mean for teams building on it.
Jul 6, 2026 · 10 min readAnthropic (Claude) Reliability: API Uptime and Status Transparency
Anthropic’s reliability posture for the Claude API - status transparency, degradation patterns, and how it compares as a production LLM dependency.
Jul 4, 2026 · 8 min readMicrosoft 365 Reliability: Exchange, Teams, and the Identity Outage Problem
Microsoft 365’s biggest outages cluster around identity and authentication. A look at the pattern, the blast radius, and the SLA fine print.
Jul 1, 2026 · 10 min readGitHub Reliability: What Its Incident History Teaches About DevOps Uptime
GitHub is a single point of failure for the world’s software supply chain. Its public incident record shows where that risk concentrates.
Jun 27, 2026 · 9 min readCoreWeave Reliability: GPU-Cloud Uptime for AI Workloads
The specialised GPU cloud behind much of the AI boom - what reliability looks like when your customers are training frontier models.
Jun 24, 2026 · 8 min readDigitalOcean Reliability: Droplet Uptime and Regional Incident Patterns
DigitalOcean’s simplicity is its selling point - but regional concentration shapes its reliability story. A look at the incident record.
Jun 20, 2026 · 8 min readAtlassian Reliability: The 2022 Outage and What It Cost Jira and Confluence Users
The two-week Atlassian outage of 2022 is a landmark case in cloud reliability - the failure mode, the recovery, and the lessons for SaaS resilience.
Jun 16, 2026 · 10 min readGitLab Reliability: Lessons From the 2017 Database Incident
GitLab’s radically transparent 2017 database-loss postmortem reshaped how the industry talks about incidents. Where its reliability stands today.
Jun 12, 2026 · 9 min read05 · Methodology & report
The dataset draws on 47 major public incidents (2017–2025), counted only where customer impact was independently observed. Full normalisation rules - what qualifies as a major incident, how duration and root cause are assigned - are on the methodology page.
Read the full narrative analysis
The State of Cloud Reliability 2026 - chaptered findings, the outlook, and the branded PDF.