Home / Companies / Neon / Blog / Post Details
Content Deep Dive

Incident Review: Pageserver outage in us-east-1

Blog post from Neon

Post Details
Company
Date Published
Author
John Spray
Word Count
1,264
Company Posts That Month
35
Language
English
Hacker News Points
-
Post removed?
No
Summary

The incident review discusses a recent outage in the us-east-1 region of Neon's services, which resulted in up to 2 hours of unavailability for approximately 0.4% of customer projects. The outage was caused by an EC2 instance failure, and it took around 30 minutes between the initial node failure and the decision to migrate projects away. The incident highlighted the need for a more resilient system, particularly in terms of fault tolerance and response time. To address this, Neon is introducing a new service called the Storage Controller, which uses a reconciliation-loop model to schedule users' Projects onto pageservers and can respond dynamically to changes in node availability or load. The Storage Controller has been in production since May 2024 and is being used to manage high-capacity Projects first, with plans to migrate all paying customers to it in the near future.

Trends Found in this Post

No tracked trend matches for this post yet.

Use This Data

Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.