Home / Companies / GitHub / Blog / Post Details
Content Deep Dive

The August 17 outage, and the work ahead

Blog post from GitHub

Post Details
Company
Date Published
Author
Vlad Fedorov
Word Count
773
Company Posts That Month
19
Language
English
Hacker News Points
643
Post removed?
No
Summary

GitHub’s August 17 outage lasted 7 hours and 47 minutes, disrupting core services including github.com, authentication, Actions, APIs, pull requests, issues, and Copilot worldwide, and followed another major Actions incident on August 6. The failure occurred when traffic reached a new peak and a critical Central US data-center component could not scale sufficiently, creating capacity pressure that cascaded across systems; recovery involved rerouting traffic, isolating infrastructure, and addressing a Copilot client retry loop that increased load. GitHub said neither incident resulted from a code or configuration change, but from insufficient capacity amid growth in monthly commits from 1.4 billion to 2.9 billion since April. In response, the company has expanded computing, storage, and network capacity, accelerated its Azure migration, improved operational testing, rollouts, monitoring, alerts, and system isolation, and introduced consistent retry limits and reviews of lower-priority resource alerts to reduce cascading failures during traffic spikes.

Trends Found in this Post
Trend Post Mentions Total Month Mentions Posts Companies MoM
AI Coding Assistant 2 1,513 470 139 -19%
Developer Experience 1 462 233 85 -22%
Observability 1 3,175 737 186 -24%
Use This Data

Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.