Home / Companies / PagerDuty / Blog / Post Details
Content Deep Dive

Outage Post-Mortem

Blog post from PagerDuty

Post Details
Company
Date Published
Author
Andrew Miklas
Word Count
1,064
Company Posts That Month
3
Language
English
Hacker News Points
-
Post removed?
No
Summary

PagerDuty experienced a 30-minute outage due to a simultaneous failure across three independent Amazon Web Services (AWS) Availability Zones in the US-East-1 region, which interrupted their ability to process and dispatch notifications. This event exposed vulnerabilities in their infrastructure and highlighted the need for greater redundancy and load management. In response, PagerDuty has implemented immediate measures, such as deploying a replica of their stack with an additional hosting provider and increasing front-end capacity, to prevent similar outages. They plan to shift away from AWS entirely and host across multiple providers to minimize correlated failure risks, and they are enhancing their load testing to handle high-traffic scenarios effectively. Additionally, PagerDuty aims to improve its customer communication during outages by introducing a Twitter account for downtime notifications and exploring phone alert systems. These efforts are part of a broader strategy to ensure reliability and regain customer trust.

Trends Found in this Post

No tracked trend matches for this post yet.

Use This Data

Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.