Home / Companies / PagerDuty / Blog / Post Details
Content Deep Dive

Injecting Failure at Netflix, Staying Reliable for 40+ Million Customers

Blog post from PagerDuty

Post Details
Company
Date Published
Author
Vivian Au
Word Count
876
Company Posts That Month
9
Language
English
Hacker News Points
-
Post removed?
No
Summary

Corey Bertram, a Site Reliability Engineer at Netflix, discussed the company's innovative approach to ensuring system reliability by deliberately injecting failures into their production systems, a strategy that has enhanced their ability to handle disruptions and maintain service for over 40 million customers. Netflix operates without a dedicated operations team, entrusting its approximately 1,000 engineers with full responsibility for their services from conception to production, fostering a culture of freedom and responsibility that encourages bold problem-solving. Due to the complexity and scale of Netflix's systems, traditional testing is impractical; instead, they automate failure testing through tools like the Simian Army to continuously challenge their systems' resilience. This approach involves focusing on cluster-level trends rather than individual incidents, automating processes, and logging every customer action to gain insights. Netflix's strategy includes promoting internal buy-in for these practices, allowing opt-outs to avoid developer burnout, and conducting failure simulations regularly, currently quarterly but moving towards bi-weekly, to ensure ongoing reliability and scalability.

Trends Found in this Post

No tracked trend matches for this post yet.

Use This Data

Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.