Open Dataset · Independent

Cloud Reliability Research & Data

The open benchmark dataset behind our numbers - AWS, Azure and Google Cloud incident frequency, duration, root cause and regional risk - plus platform-by-platform reliability insights. The written analysis lives in the 2026 report.

Corpus
47 major incidents
Coverage
AWS · Azure · GCP
Period
2017–2025
Last updated
June 30, 2026

01 · Benchmark data

The open corpus behind our figures: 47 major public incidents across AWS, Azure and Google Cloud, 2017–2025, normalised per our severity methodology. Every chart sits next to a text summary so the numbers stay quotable.

47

major incidents analysed, 2017–2025

3.4h

median duration of a major incident

11h

90th-percentile duration

Major public incidents by provider (2017–2025)
AWS19 major incidents
Microsoft Azure16 major incidents
Google Cloud12 major incidents
Median incident duration by provider
AWS3h 37m median
Microsoft Azure3h 46m median
Google Cloud2h 48m median
Root causes of major incidents (share)
Configuration / change error45%
Software defect21%
Capacity / overload13%
Network / hardware failure11%
External attack (DDoS)6%
Other / undisclosed4%
Flagship-region & global-layer concentration
AWS incidents involving us-east-163%
Azure incidents involving global services44%
GCP incidents involving global layers50%
Major incidents per year (all providers)
2019incidents / year2025

02 · Reliability positioning

Each provider plotted on demonstrated reliability against operational maturity; bubble size reflects market presence. All three qualify as Leaders or Strong Performers - the gaps are in execution, not capability.

GCPAWSAzureOperational maturity →Demonstrated reliability →
  • GCP - LeadersReliability 89 · Maturity 82 · Presence 62
  • AWS - LeadersReliability 83 · Maturity 88 · Presence 95
  • Azure - Strong PerformersReliability 66 · Maturity 71 · Presence 84

03 · Where downtime comes from

Across all three providers, networking and DNS remained the single largest root-cause category - a pattern that has held for three years running. Software deploys are the fastest-shrinking category as progressive rollout and automated rollback mature.

  • Networking / DNS31%
  • Power & Hardware22%
  • Software Deploys19%
  • Capacity / Scaling16%
  • Configuration12%

04 · Research insights

Reliability profiles of the platforms teams depend on most - each grounded in public incident records and reporting from CRN, TechTarget, the Futurum Group and Time.

Global edge network of interconnected points of light spanning the Earth at nightEdge & CDN

Cloudflare Reliability: The Edge Dependency Behind Half the Web

Cloudflare sits in front of a huge share of the internet, so its rare control-plane and config outages have an outsized blast radius. A look at the pattern and how to design for it.

Jul 16, 2026 · 9 min read
An analytics dashboard on a laptopObservability

Datadog Reliability: Why Your Monitoring Must Not Share Fate With Your App

Observability platforms fail too, as the March 2023 multi-region Datadog outage showed. When your monitoring goes dark at the same time as your app, you are flying blind.

Jul 14, 2026 · 9 min read
Credit card and online payment representing a checkout being processedPayments

Stripe Reliability: Designing Checkout for Graceful Degradation

Stripe is the payment backbone for a large share of online commerce, so a payments incident can stop revenue cold. How to keep checkout resilient when the payment layer degrades.

Jul 11, 2026 · 9 min read
Team collaborating around laptops with real-time chat tools openPlatform Reliability

Slack Reliability: What a Decade of Outages Teaches About Real-Time Uptime

Slack’s incident history - from the 2021 new-year outage to routine degradations - and what it reveals about running always-on collaboration at scale.

Jul 8, 2026 · 9 min read
Abstract AI and machine-learning visualization representing the OpenAI APIAI Infrastructure

OpenAI API Reliability: Uptime, Capacity, and the Incident Patterns Behind ChatGPT

How OpenAI’s API and ChatGPT have held up under explosive demand - the recurring capacity and degradation patterns, and what they mean for teams building on it.

Jul 6, 2026 · 10 min read
Abstract AI and machine-learning visualization representing the Claude APIAI Infrastructure

Anthropic (Claude) Reliability: API Uptime and Status Transparency

Anthropic’s reliability posture for the Claude API - status transparency, degradation patterns, and how it compares as a production LLM dependency.

Jul 4, 2026 · 8 min read
Team collaborating around laptops on Microsoft 365 appsEnterprise SaaS

Microsoft 365 Reliability: Exchange, Teams, and the Identity Outage Problem

Microsoft 365’s biggest outages cluster around identity and authentication. A look at the pattern, the blast radius, and the SLA fine print.

Jul 1, 2026 · 10 min read
Source code on a dark screenDeveloper Platforms

GitHub Reliability: What Its Incident History Teaches About DevOps Uptime

GitHub is a single point of failure for the world’s software supply chain. Its public incident record shows where that risk concentrates.

Jun 27, 2026 · 9 min read
High-performance graphics cardsAI Infrastructure

CoreWeave Reliability: GPU-Cloud Uptime for AI Workloads

The specialised GPU cloud behind much of the AI boom - what reliability looks like when your customers are training frontier models.

Jun 24, 2026 · 8 min read
A data center corridor lined with cablingCloud Infrastructure

DigitalOcean Reliability: Droplet Uptime and Regional Incident Patterns

DigitalOcean’s simplicity is its selling point - but regional concentration shapes its reliability story. A look at the incident record.

Jun 20, 2026 · 8 min read
Team meeting in an office tracking issues across Jira and ConfluenceEnterprise SaaS

Atlassian Reliability: The 2022 Outage and What It Cost Jira and Confluence Users

The two-week Atlassian outage of 2022 is a landmark case in cloud reliability - the failure mode, the recovery, and the lessons for SaaS resilience.

Jun 16, 2026 · 10 min read
Colorful source code on a screenDeveloper Platforms

GitLab Reliability: Lessons From the 2017 Database Incident

GitLab’s radically transparent 2017 database-loss postmortem reshaped how the industry talks about incidents. Where its reliability stands today.

Jun 12, 2026 · 9 min read

05 · Methodology & report

The dataset draws on 47 major public incidents (2017–2025), counted only where customer impact was independently observed. Full normalisation rules - what qualifies as a major incident, how duration and root cause are assigned - are on the methodology page.

Read the full narrative analysis

The State of Cloud Reliability 2026 - chaptered findings, the outlook, and the branded PDF.

Read the 2026 report