Home / Companies / Cockroach Labs / Blog / Post Details
Content Deep Dive

Santander: Hacking human error to achieve operational resilience

Blog post from Cockroach Labs

Post Details
Company
Date Published
Author
Michelle Gienow
Word Count
1,061
Company Posts That Month
12
Language
English
Hacker News Points
-
Post removed?
No
Summary

CockroachDB: The Definitive Guide emphasizes the importance of architecting applications for scalability, resilience, and low-latency performance, especially in the face of disasters such as fires and cloud provider outages. At RoachFest23, Thomas Boltze from Santander highlighted that human error is often the primary cause of system failures, indicating the need for robust resiliency practices. He shared insights from Santander's journey to achieving a resilient payments system by continuously testing, identifying, and addressing failures, which eventually enabled the system to withstand data center outages and process payments without interruption. The key to their success lay in a culture shift towards curiosity, shared responsibility, and automation, resulting in a system designed to handle multi-region and multi-cloud failures, ensuring uninterrupted service even during significant outages.

Trends Found in this Post
Trend Post Mentions Total Month Mentions Posts Companies MoM
Observability 1 1,229 243 85 -6%
Use This Data

Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.