Incident post-mortem analysis: Partial unavailability on November 26, 2024
Blog post from Nebius
A faulty release in the VM recovery sequence in the eu-north1 region led to 282 virtual machines (VMs) restarting, with 164 experiencing extended downtime that required manual intervention, resulting in customer workload interruptions. The incident was triggered by an issue in the Compute API service, which mistakenly assumed certain VMs were non-operational due to flawed event processing assumptions and lack of optimization for recovery operations, causing some VMs to become stuck. The response involved halting automatic recovery operations, categorizing affected VMs, and applying thoroughly tested mitigation procedures, which included stopping and restarting VMs in controlled batches. The incident underscored the need for improved automation in recovery procedures to enhance response times, and a post-incident action plan was developed focusing on pre-deployment validation, operational improvements, and enhanced communication protocols to prevent future occurrences.
No tracked trend matches for this post yet.
Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.