01 / 01

Understanding one of GitHub's most significant service disruptions

The August 17 Outage

Incident Analysis

August 2023

02

What Happened

A timeline of the August 17 incident

Incident Timeline

08:00 UTC
Initial Detection

Systems began showing signs of degradation across GitHub services

08:45 UTC
Widespread Impact

Multiple services became unavailable, affecting millions of users globally

10:30 UTC
Root Cause Identified

Database infrastructure failure traced to configuration change

15:30 UTC
Full Recovery

Services fully restored after extensive recovery operations

7.5 hours
Total duration of the outage
Millions of developers worldwide impacted
05

Root Cause & Response

Understanding the technical failure and how it was resolved

What Went Wrong

  • 01
    Database Configuration Change A maintenance operation introduced unexpected behavior in MySQL cluster
  • 02
    Cascading Failures Primary database failure triggered automatic failover, but secondary systems couldn't handle the load
  • 03
    Infrastructure Interdependencies Multiple GitHub services depend on the same database layer, amplifying impact

Recovery Process

1

Isolate Impact

Contained the issue to prevent further spread

2

Restore from Backups

Initiated recovery procedures using redundant systems

3

Gradual Re-enablement

Carefully brought services back online in phases

4

Verification & Testing

Confirmed system stability before full restoration

Key Takeaways

Building more resilient systems through learning

github.blog/news-insights/company-news/the-august-17-outage
Made with AirSlide
𝕏 in