March 2021 Summaries
3 posts from Gremlin
Filter
Month:
Year:
Post Summaries
Back to Blog
Jérôme Petazzoni, a container technology educator and former Docker engineer, shares insights on chaos engineering and system reliability in an episode of the "Break Things on Purpose" podcast. He recounts a significant incident from his time at dotCloud, the precursor to Docker, where a misunderstanding of Riak's replication mechanism led to a critical data loss, highlighting the importance of defensive design and recovery strategies in complex systems. Petazzoni emphasizes the necessity of designing applications to recover gracefully from failures rather than preventing failures entirely, drawing parallels to Kubernetes and its layered, modular architecture. He advocates for a focus on recovery strategies in complex systems and emphasizes the importance of enabling technologies that support essential work in fields such as education and healthcare. Petazzoni concludes by encouraging individuals and organizations to adopt a mindset that anticipates failure and prioritizes resilience, which he believes is crucial for both technological progress and societal benefits.
Mar 23, 2021
4,579 words in the original blog post.
In a podcast episode of "Break Things on Purpose," J Paul Reed, Senior Applied Resilience Engineer at Netflix, discusses his unique role in enhancing the reliability of Netflix's systems through resilience engineering. Reed, who is part of Netflix's CORE team (Critical Operations and Reliability Engineering), explains that his responsibilities include conducting incident reviews, analyzing socio-technical risk patterns, and facilitating discussions that encourage teams to identify and address risks before they escalate into incidents. He emphasizes the importance of understanding both technical and social aspects of systems, as well as creating spaces for emergent discussions to foster organizational learning and resilience. Reed also highlights the shift from traditional Newtonian thinking to Quantum thinking in resilience engineering, where understanding the complex interactions within socio-technical systems is crucial. The conversation touches on the challenges of adapting to remote work during the COVID-19 pandemic and the importance of creating environments where engineers feel comfortable discussing potential risks. Overall, the episode provides insights into the evolving field of resilience engineering and its critical role in maintaining the reliability of complex systems like Netflix.
Mar 09, 2021
6,313 words in the original blog post.
API gateways are crucial components in cloud-native systems, performing functions like request routing, authentication, caching, and more. Failures in these gateways can jeopardize entire deployments, making their resilience a priority. The text explores the application of Chaos Engineering to test and enhance the resilience of API gateways by simulating conditions such as backend outages and high network latency. It highlights the use of tools like Gremlin to conduct controlled experiments, validate load balancing, caching mechanisms, and the general reliability of the gateway itself. Examples include simulating service unavailability to test load balancing, using latency attacks to assess caching effectiveness, and conducting shutdown attacks to test Kubernetes' ability to restart failed pods. The text suggests building a practice of frequent automated experiments to maintain API reliability as system configurations evolve.
Mar 04, 2021
1,984 words in the original blog post.