Home / Companies / Railway / Blog / Post Details
Content Deep Dive

Incident Report: Dec 13th, 2023

Blog post from Railway

Post Details
Company
Date Published
Author
Angelo Saraceno
Word Count
834
Company Posts That Month
3
Language
English
Hacker News Points
-
Post removed?
No
Summary

The Railway engineering team experienced a production outage from 21:45 UTC to 23:24 UTC due to issues with Google's Metadata server, which was upgraded as part of the rollout of GKE v1.25. The upgrade caused significant delays in encrypt and decrypt requests, affecting the ability of users to deploy new workloads and update environment variables. The team quickly identified and addressed the issue by rolling out a newer version of GKE that was not affected by the known issue, and restored all services to normal operation by 23:24 UTC. The incident highlighted the importance of careful planning and monitoring during significant infrastructure changes, and led to several takeaways for improving internal coordination, messaging, re-shoring, and incident management.

Trends Found in this Post
Trend Post Mentions Total Month Mentions Posts Companies MoM
Kubernetes 3 1,238 142 66 -27%
Platform Engineering 1 316 54 29 -24%
Secrets Management 1 369 78 52 -42%
Use This Data

Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.