Home / Companies / Neon / Blog / Post Details
Content Deep Dive

How we systematically improved our reliability

Blog post from Neon

Post Details
Company
Date Published
Author
Dmitrii Mokhnatkin
Word Count
2,450
Company Posts That Month
15
Language
English
Hacker News Points
-
Post removed?
No
Summary

Neon describes a reliability program for Lakebase Postgres, whose control plane manages millions of databases across Neon and Databricks, three cloud providers, more than 30 regions, and over 70 isolated deployment cells. After two incidents triggered by external dependency failures created self-reinforcing control-plane outages, the company adopted more rigorous postmortems focused on measured customer impact, detection and mitigation times, evidence-based root-cause analysis, and specific, owned remediation items. Resulting work included more than 30 fixes involving the control plane’s Postgres queries and transaction handling, inherited library and network defaults, queue processing, backpressure, and fairness under load. Neon also introduced a preventive framework that scores changes by probability and impact, applying proportionate reliability checks during design, development, release, and operation, such as premortems, load and failure testing, staged rollouts, SLO monitoring, rollback validation, and on-call drills. Future efforts include failure-mode and fault-tree analysis plus automated stress, limit, and recovery testing to identify weaknesses before customers are affected.

Trends Found in this Post
Trend Post Mentions Total Month Mentions Posts Companies MoM
AI Agents 1 931 231 103 -84%
Observability 1 472 102 54 -85%
Real-time 1 649 155 80 -85%
Use This Data

Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.