Home / Companies / Nebius / Blog / Post Details
Content Deep Dive

Incident post-mortem analysis: us-central1 service disruption on September 3, 2025

Blog post from Nebius

Post Details
Company
Date Published
Author
Nebius team
Word Count
912
Company Posts That Month
13
Language
English
Hacker News Points
-
Post removed?
No
Summary

On September 3, 2025, Nebius experienced a 1 hour and 45-minute service disruption in the us-central1 region due to a routing configuration conflict between network domains, which was triggered by a combination of a past routing policy optimization and a recent security infrastructure extension. This conflict led to a persistent routing loop that affected multiple customer-facing services and exposed cross-regional dependencies, amplifying the impact beyond the immediate region. Although other regions remained functional, inconsistent service experiences were reported as the load balancing system continued directing users to both functional and non-functional endpoints. The disruption impacted public API operations, console communication, virtual machine management, tenant registration, and developer tools targeting the affected region, while those targeting other regions experienced limited resource visibility. The root cause was traced to a latent configuration issue that was not detected during testing and a subsequent routing recomputation that perpetuated the loop. In response, Nebius has outlined improvements in network infrastructure and service architecture, such as implementing stricter routing controls, enhancing monitoring and alerting systems, reducing cross-regional dependencies, and improving failover mechanisms to prevent future incidents and minimize customer impact.

Trends Found in this Post

No tracked trend matches for this post yet.

Use This Data

Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.