Home / Companies / Nebius / Blog / Post Details
Content Deep Dive

Incident post-mortem analysis: us-central1 service disruption on August 19, 2026

Blog post from Nebius

Post Details
Company
Date Published
Author
Nebius team
Word Count
2,354
Company Posts That Month
6
Language
English
Hacker News Points
-
Post removed?
No
Summary

A storm-related failure at the us-central1 data center on August 19, 2026 disabled both the building management system and chilled-water cooling, causing temperatures to rise rapidly and triggering thermal shutdowns across servers, network equipment, and rack power systems. The resulting regional outage disrupted compute, VM networking, storage, Managed Kubernetes, Token Factory, managed platform services, and regional console and API endpoints, while other regions remained largely unaffected. Cooling was restored and the facility stabilized within about six hours, but recovery continued into August 20 because the regional control plane required manual restoration, VM recovery mechanisms were rate-limited and did not automatically retry failures, stale disk attachments blocked some resources, and a disk hot-plug bug left Kubernetes control-plane data disks unattached. Object storage recovered by 13:27 UTC, external VM connectivity by 12:50 UTC, Kubernetes clusters by 21:58 UTC, Token Factory by 20:25 UTC, and reserved GPU capacity by 08:38 UTC the following day, although some on-demand capacity remained reduced during hardware repair. The provider identified gaps in independent facility monitoring, environmental alerting, full-region recovery procedures, automation for mass failures, regional observability, and traffic health checks, and plans resilience reviews, independent alerts, emergency procedures, regular recovery drills, improved storage reconciliation, and cross-region monitoring.

Trends Found in this Post

No tracked trend matches for this post yet.

Use This Data

Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.