Home / Companies / Nebius / Blog / Post Details
Content Deep Dive

Incident post-mortem analysis: outage of the S3 service in the eu-north1 region

Blog post from Nebius

Post Details
Company
Date Published
Author
Nebius team
Word Count
840
Company Posts That Month
9
Language
English
Hacker News Points
-
Post removed?
No
Summary

On May 5, 2025, an outage occurred in the S3 service within the eu-north1 region due to increased migration traffic and unexpected program behavior, leading to CPU resource exhaustion in the YDB database's thread pool. This caused a buffer overflow in the storage group, which was misinterpreted as an 'out of space' issue, misleading the SRE team and rendering the database inoperable. The incident resulted in a total service unavailability for clients using S3 object storage in the region from 15:30 to 17:00 UTC, although no data was lost. The root cause was traced to a flaw in the thread pool's automatic configuration system, which mismanaged CPU allocation due to its measurement approach. The issue was compounded by the absence of backpressure on buffer writing and the presence of long-running tasks in the batch pool, which blocked short-running tasks. Recovery involved manual traffic rerouting and database reconfiguration, with service fully restored by 17:50 UTC. To prevent future occurrences, corrective measures include disabling automatic pool size configuration, enhancing alert systems, and implementing code changes to better manage resource allocation and task execution.

Trends Found in this Post

No tracked trend matches for this post yet.

Use This Data

Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.