Home / Companies / Gremlin / Blog / February 2022

February 2022 Summaries

3 posts from Gremlin

Filter
Month: Year:
Post Summaries Back to Blog
On December 7, 2021, Amazon Web Services (AWS) faced a significant outage in its US-East-1 region, which highlighted the complex interplay of systems within cloud infrastructure and the importance of reliability engineering. The incident, triggered by an automated scaling event, led to unexpected behavior and network congestion impacting some AWS services while sparing others. Factors such as impaired monitoring, affected deployment systems, and the need for careful remediation to avoid further disruptions complicated the resolution process. The outage underscored the necessity of effective monitoring, rapid deployment of fixes, and adherence to safe deployment practices, such as canary or staggered strategies. Additionally, the Synchronization of Chaos theory suggests that integrating more AWS services can mitigate chaos, and AWS’s Well-Architected Framework offers guidance on using availability zones and multi-region designs for enhanced reliability. To mitigate similar risks, practices like validating monitoring tools through controlled incidents, ensuring backup systems for deployments, and leveraging redundancies are essential.
Feb 22, 2022 1,293 words in the original blog post.
Carissa Morrow's journey into the tech industry exemplifies resilience and adaptability, transitioning from a career as a certified ophthalmic technician to becoming a cloud engineer at ClickBank in just three years. Her story, shared on the "Break Things on Purpose" podcast, highlights the challenges and learning experiences she encountered, including breaking production systems and the importance of asking questions and seeking mentorship. Morrow emphasizes the value of resilience and continual learning in technology, as well as the necessity for organizations to avoid complacency, which can lead to what she refers to as the "organizational death spiral." Her experiences underscore the need for effective communication, problem-solving, and the strategic use of monitoring and testing tools to enhance system resilience and reliability.
Feb 22, 2022 5,275 words in the original blog post.
Gunnar Grosch, a Senior Developer Advocate at AWS, discusses his journey from being an AWS user to becoming an AWS Serverless Hero and his current role in promoting serverless reliability through Chaos Engineering. He emphasizes the importance of testing serverless systems for resilience, despite their inherent reliability, by using techniques like Chaos Engineering to examine how different serverless components interact under stress. Grosch highlights the challenges of maintaining documentation and mental models in dynamic serverless environments and advocates for integrating observability and tracing early in the development process. He also introduces AWS Resilience Hub, a new service that leverages the AWS Well-Architected Framework to help users assess and improve their system's resilience, offering recommendations for enhancing architecture and running chaos experiments through AWS Fault Injection Simulator. Grosch encourages hands-on experience with these tools to promote the adoption of Chaos Engineering practices.
Feb 08, 2022 4,931 words in the original blog post.