Home / Companies / Nebius / Blog / Post Details
Content Deep Dive

Incident post-mortem analysis: us-central1 service disruption on March 10, 2026

Blog post from Nebius

Post Details
Company
Date Published
Author
Nebius team
Word Count
1,020
Company Posts That Month
16
Language
English
Hacker News Points
-
Post removed?
No
Summary

On March 10, 2026, a scheduled power infrastructure maintenance by a provider unexpectedly escalated into a broader data center power incident, affecting multiple services in the us-central1 region. Initially planned to have minimal customer impact, the maintenance resulted in several unplanned power failures, leading to a series of disruptions including loss of external VM connectivity, inaccessibility of Managed Kubernetes clusters, and interruptions in public S3 endpoints. The incident unfolded in multiple stages over several hours, with repeated power interruptions complicating recovery efforts. The root cause was linked to the maintenance's deviation from the planned scope, involving incorrect switching and failure of rack ATS units during a temporary power configuration. This incident highlighted the need for improved communication and planning with the power provider, more conservative preparation standards, and the development of a counter-emergency recovery procedure to ensure faster recovery in future incidents. The post-incident action plan includes revising maintenance communication standards, enhancing planning quality for future electrical work, and conducting regular recovery drills to ensure a more efficient and predictable restoration process.

Trends Found in this Post

No tracked trend matches for this post yet.

Use This Data

Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.