How we systematically improved our reliability
Blog post from Neon
Neon describes a reliability program for Lakebase Postgres, whose control plane manages millions of databases across Neon and Databricks, three cloud providers, more than 30 regions, and over 70 isolated deployment cells. After two incidents triggered by external dependency failures created self-reinforcing control-plane outages, the company adopted more rigorous postmortems focused on measured customer impact, detection and mitigation times, evidence-based root-cause analysis, and specific, owned remediation items. Resulting work included more than 30 fixes involving the control plane’s Postgres queries and transaction handling, inherited library and network defaults, queue processing, backpressure, and fairness under load. Neon also introduced a preventive framework that scores changes by probability and impact, applying proportionate reliability checks during design, development, release, and operation, such as premortems, load and failure testing, staged rollouts, SLO monitoring, rollback validation, and on-call drills. Future efforts include failure-mode and fault-tree analysis plus automated stress, limit, and recovery testing to identify weaknesses before customers are affected.
| Trend | Post Mentions | Total Month Mentions | Posts | Companies | MoM |
|---|---|---|---|---|---|
| AI Agents | 1 | 931 | 231 | 103 | -84% |
| Observability | 1 | 472 | 102 | 54 | -85% |
| Real-time | 1 | 649 | 155 | 80 | -85% |
Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.