November 2018 Summaries
3 posts from Gremlin
Filter
Month:
Year:
Post Summaries
Back to Blog
Mikolaj Pawlikowski, speaking at Chaos Conf 2018, detailed Bloomberg's experience with Chaos Engineering, emphasizing the challenges of managing distributed systems on their DTP microservices platform. To tackle these challenges, his team developed PowerfulSeal, a tool designed to introduce controlled chaos into Kubernetes environments. Unlike Chaos Monkey, PowerfulSeal offers greater customization and operates in four modes: interactive, label, demo, and autonomous, each providing various levels of control and automation over system disturbances. The tool is aimed at increasing confidence in system reliability by simulating failures and identifying weaknesses, rather than proving system robustness outright. Pawlikowski highlighted the importance of convincing management of the necessity of such practices, often citing the inevitability of system failures due to scale. PowerfulSeal, open-sourced at Kubecon, invites user feedback to improve its utility in identifying and addressing system vulnerabilities.
Nov 16, 2018
2,884 words in the original blog post.
Vilas Veeraraghaven's presentation at Chaos Conf 2018 focused on transitioning from chaos engineering to resilience engineering at Walmart, emphasizing the importance of preparing systems for resilience in the face of chaos rather than just introducing chaos. Drawing from his experience at Netflix, Veeraraghaven shared how Walmart has been implementing a structured approach to enhance system resilience, involving multiple levels that progressively integrate automation and testing to reduce support costs and revenue losses. The initiative encourages teams to develop disaster recovery playbooks, conduct failure injection tests, and use tools like Gremlin to simulate failures and refine their responses. The effort has fostered a culture of accountability and collaboration within Walmart, leading to the establishment of a community of chaos practitioners who share best practices and support each other in achieving higher levels of resilience. Despite some challenges, such as a significant outage due to a storm in Texas, the approach has empowered teams, reduced silos, and improved overall system resilience, with plans to further automate and potentially open-source the developed tools.
Nov 09, 2018
2,490 words in the original blog post.
Adrian Cockroft's keynote at Chaos Conf 2018 focuses on Chaos Engineering, emphasizing its role in preparing systems to handle failures effectively. He discusses the evolution of Chaos Engineering from traditional disaster recovery practices, highlighting its importance in today's cloud-based infrastructure where systems must be resilient to a range of failures, from hardware malfunctions to software bugs and operational missteps. Cockroft emphasizes the need for a culture that encourages reporting and learning from small incidents to prevent larger failures, advocating for continuous automated testing rather than annual disaster recovery exercises. He also highlights the importance of observability and the role of human judgment in managing unforeseen failures, underscoring the integration of Chaos Engineering into organizational processes to enhance the reliability and safety of complex systems.
Nov 02, 2018
10,647 words in the original blog post.