Home / Companies / PagerDuty / Blog / Post Details
Content Deep Dive

Surviving a Datacenter Outage

Blog post from PagerDuty

Post Details
Company
Date Published
Author
Amanda Folson
Word Count
556
Company Posts That Month
4
Language
English
Hacker News Points
-
Post removed?
No
Summary

PagerDuty is designed to ensure timely notifications reach the right individuals during system failures by leveraging a robust alerting pipeline that utilizes distributed technologies for redundancy and fault tolerance. The pipeline begins with an event endpoint and involves several services, culminating in a messaging service that alerts people, relying on technologies like Scala, Cassandra, and Zookeeper. Cassandra has been used in production for over two years, with a relatively small dataset that is purged after events are resolved, and is configured in clusters across multiple datacenters to ensure durability and availability through a quorum consistency level. The cluster design, though atypical due to inter-datacenter latency, prioritizes durability and prevents data loss even if a datacenter fails, and the system is rigorously tested with intentional failures to ensure resilience. PagerDuty is actively seeking to enhance its infrastructure and is hiring for positions in San Francisco and Toronto.

Trends Found in this Post

No tracked trend matches for this post yet.

Use This Data

Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.