Home / Companies / Stytch / Blog / Post Details
Content Deep Dive

Stytch postmortem 2023-02-23

Blog post from Stytch

Post Details
Company
Date Published
Author
Ovadia Harary
Word Count
2,240
Company Posts That Month
7
Language
English
Hacker News Points
-
Post removed?
No
Summary

On February 23, 2023, Stytch experienced a full system outage due to an infrastructure configuration change that inadvertently deleted an instance profile, affecting their Kubernetes worker nodes and resulting in downtime for their Live API, Frontend SDKs, and Dashboard. The outage was traced back to the removal of managed node groups and the subsequent deletion of an instance profile that was critical for Karpenter, Stytch's dynamic node provisioning tool. This incident highlighted a gap in the AWS documentation regarding the cascading effects of node group deletions, which led to the misconfiguration of IAM roles. To address this, Stytch implemented several action items, including improving alert severity, separating cloud resources, and overhauling their EKS and Karpenter configurations. The company also engaged with AWS to better understand the undocumented actions and committed to enhancing their incident response processes to prevent similar occurrences in the future.

Trends Found in this Post

No tracked trend matches for this post yet.

Use This Data

Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.