Home / Companies / Cloudflare / Blog / Post Details
Content Deep Dive

A Byzantine failure in the real world

Blog post from Cloudflare

Post Details
Company
Date Published
Author
Tom Lianza, Chris Snook
Word Count
1,913
Company Posts That Month
25
Language
English
Hacker News Points
16
Post removed?
No
Summary

On November 2, 2020, Cloudflare experienced an incident that impacted the availability of its API and dashboard for six hours and 33 minutes. The issue was caused by a Byzantine fault, which led to a cascading series of events involving partial switch failure, etcd errors, promotion of new primary databases, and overloaded authentication databases. Despite having redundancy in each system, the combination of degraded states made it difficult to model and anticipate the chain of events that transpired. The incident led Cloudflare to revisit its configuration parameters for auto-remediation processes and prompted further research into Byzantine Fault Tolerance (BFT) consensus protocols.

Trends Found in this Post
Trend Post Mentions Total Month Mentions Posts Companies MoM
Real-time 1 786 208 71 -4%
Use This Data

Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.