Home / Companies / Datadog / Blog / Post Details
Content Deep Dive

2023-03-08 Incident: A Deep Dive into the Platform-level Recovery

Blog post from Datadog

Post Details
Company
Date Published
Author
Laurent Bernaille
Word Count
4,508
Company Posts That Month
33
Language
English
Hacker News Points
-
Post removed?
No
Summary

On March 8, 2023, Datadog experienced a global outage that affected all services across multiple regions due to an unexpected system patch. The company's teams worked for approximately 13 hours to restore the majority of their compute capacity across all regions. They faced several challenges during this process, including different responses required for each region, limitations on the maximum number of VM instances in a peering group, and reaching subnet capacity limits. Despite these obstacles, Datadog managed to recover its platform-level capabilities and continue working towards complete recovery.

Trends Found in this Post
Trend Post Mentions Total Month Mentions Posts Companies MoM
Kubernetes 21 1,603 204 73 -5%
Secrets Management 19 1,263 106 61 +23%
Observability 2 1,433 240 77 -8%
Use This Data

Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.