Home / Companies / Nebius / Blog / Post Details
Content Deep Dive

Incident post-mortem analysis: Major connectivity loss on January 27, 2025

Blog post from Nebius

Post Details
Company
Date Published
Author
-
Word Count
916
Company Posts That Month
9
Language
English
Hacker News Points
-
Post removed?
No
Summary

On January 27, 2025, a routine release in the Compute and VPC control plane cluster in the eu-north1 region led to a cascading failure, severely affecting core infrastructure services, including a complete failure of compute API operations and the loss of external connectivity for user virtual machines. The incident was triggered by a sudden spike in API requests combined with misconfigurations in service degradation mechanisms, resulting in uncontrolled resource consumption and service outages. The incident response team isolated control plane nodes to regain control, restored administrative access, and implemented rate-limiting mechanisms to stabilize the system. By 23:12 UTC, all services were fully operational, and temporary mitigations were reviewed. The root cause was identified as a combination of a significant spike in requests due to a routine release and misconfigured service degradation subsystems. An action plan was developed to enhance control plane stability, improve degradation frameworks, increase network resilience, and refine operational processes to prevent recurrences and improve system reliability.

Trends Found in this Post

No tracked trend matches for this post yet.

Use This Data

Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.