Stytch postmortem 2023-02-23
Blog post from Stytch
On February 23, 2023, Stytch experienced a full system outage due to an infrastructure configuration change that inadvertently deleted an instance profile, affecting their Kubernetes worker nodes and resulting in downtime for their Live API, Frontend SDKs, and Dashboard. The outage was traced back to the removal of managed node groups and the subsequent deletion of an instance profile that was critical for Karpenter, Stytch's dynamic node provisioning tool. This incident highlighted a gap in the AWS documentation regarding the cascading effects of node group deletions, which led to the misconfiguration of IAM roles. To address this, Stytch implemented several action items, including improving alert severity, separating cloud resources, and overhauling their EKS and Karpenter configurations. The company also engaged with AWS to better understand the undocumented actions and committed to enhancing their incident response processes to prevent similar occurrences in the future.
No tracked trend matches for this post yet.
Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.