Home / Companies / Nebius / Blog / Post Details
Content Deep Dive

Incident post-mortem analysis: networking issues for Managed Services for Kubernetes on March 13, 2025

Blog post from Nebius

Post Details
Company
Date Published
Author
-
Word Count
856
Company Posts That Month
11
Language
English
Hacker News Points
-
Post removed?
No
Summary

On March 13, 2025, a release of a control plane component in the VPC service's eu-west1 region led to significant networking disruptions for Managed Services for Kubernetes, due to a bug in the network configuration data processing that affected the IP alias mechanism. Although the bug had existed in earlier versions, it was triggered by a new feature in the latest release, complicating efforts to mitigate the incident through rollback procedures. From 13:36 to 18:11 UTC, nearly all pods in the eu-west1 region experienced connectivity loss, while the eu-north1 region remained unaffected. The incident was traced to a bug in the VPC control plane's handling of BGP messages, where route distinguishers were partially ignored, leading to incorrect merging of IP alias announcements. This issue was difficult to detect in testing and canary deployments due to specific conditions required for its manifestation. The response plan includes improving service observability, enhancing incident response capabilities, refining testing and deployment procedures, and introducing alerts on critical network metrics.

Trends Found in this Post

No tracked trend matches for this post yet.

Use This Data

Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.