Incident post-mortem analysis: us-central1 service disruption on March 10, 2026
Blog post from Nebius
On March 10, 2026, a scheduled power infrastructure maintenance by a provider unexpectedly escalated into a broader data center power incident, affecting multiple services in the us-central1 region. Initially planned to have minimal customer impact, the maintenance resulted in several unplanned power failures, leading to a series of disruptions including loss of external VM connectivity, inaccessibility of Managed Kubernetes clusters, and interruptions in public S3 endpoints. The incident unfolded in multiple stages over several hours, with repeated power interruptions complicating recovery efforts. The root cause was linked to the maintenance's deviation from the planned scope, involving incorrect switching and failure of rack ATS units during a temporary power configuration. This incident highlighted the need for improved communication and planning with the power provider, more conservative preparation standards, and the development of a counter-emergency recovery procedure to ensure faster recovery in future incidents. The post-incident action plan includes revising maintenance communication standards, enhancing planning quality for future electrical work, and conducting regular recovery drills to ensure a more efficient and predictable restoration process.
No tracked trend matches for this post yet.
Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.