Home / Companies / Gremlin / Blog / May 2020

May 2020 Summaries

5 posts from Gremlin

Filter
Month: Year:
Post Summaries Back to Blog
Andre Newman, a technical writer at Gremlin, has quickly made a significant impact in his role by authoring articles on Chaos Engineering and related topics. Before joining Gremlin, Andre transitioned from development work at LabTech to freelance writing, focusing on DevOps and cloud computing, where he contributed to companies like Loggly. He emphasizes the importance of networking for freelancers and shares that building a portfolio through freelance exchanges was key to his success. Andre joined Gremlin because it offered the opportunity to explore the emerging field of Chaos Engineering while maintaining the benefits of a remote work lifestyle. Outside of his professional life, Andre enjoys writing fiction, engaging in outdoor activities, experimenting with Arduino projects, and volunteering with local 3D printing groups to produce face shields during the COVID-19 pandemic.
May 24, 2020 1,120 words in the original blog post.
Amazon DynamoDB is a highly available and durable NoSQL database service that offers features like automatic replication, backups, and optional multi-region replication, making it suitable for large-scale, near real-time data access. Despite its strengths, DynamoDB deployments can face reliability challenges due to various factors such as integration with other services, application implementation issues, and occasional service disruptions. Chaos Engineering is recommended as a method to identify and mitigate these risks by systematically testing and verifying the reliability of applications using DynamoDB. This approach involves conducting experiments to simulate failure scenarios in order to enhance application resilience and ensure a seamless user experience. While DynamoDB provides a robust platform for building reliable databases, understanding and addressing potential vulnerabilities is crucial for maintaining performance and availability.
May 21, 2020 1,520 words in the original blog post.
Gremlin has announced the release of its chaos engineering agent for Windows, enabling engineers to conduct resilience testing on Windows systems, from Server 2008 R2 and later to Windows 7 and later. This development is significant as Windows servers account for a large portion of the server market, necessitating rigorous testing of Windows applications for reliability. The Gremlin Windows agent allows users to perform various chaos experiments, such as Shutdown, CPU, Disk, I/O, Memory, Blackhole, and Latency attacks, with more types of attacks in development. The agent is designed to be user-friendly, facilitating the installation, configuration, and visualization of Windows infrastructure through Gremlin's web app or API. By tagging and customizing identifiers for each machine, users can strategically target or randomly select systems for testing. The goal of these experiments is to ensure that critical Microsoft enterprise applications, such as Windows Server Failover clustering and SQL Server replication, can withstand real-world disruptions, thus enhancing the reliability of business-critical applications. Gremlin offers a 30-day free trial for users interested in exploring the platform's capabilities.
May 13, 2020 645 words in the original blog post.
As remote work causes unprecedented network stress, companies are adapting to increased traffic by expanding cloud services and implementing robust reliability strategies. With video calls contributing significantly to the surge, organizations must ensure their systems, often not designed for such heavy external use, remain reliable. To maintain productivity, businesses are enhancing load balancing, utilizing additional cloud instances, and implementing continuity plans. Strategies to handle this include solidifying on-call schedules, creating concise system one-pagers, reviewing past incident retrospectives, and engaging in chaos engineering to test system resilience. Early and frequent preparation is crucial, encompassing measures like expanding compute nodes, establishing failover methods, and implementing automated redundancy. By doing so, companies aim to mitigate potential disruptions and maintain service availability during traffic spikes, ensuring a seamless remote working environment.
May 12, 2020 1,306 words in the original blog post.
Failover Conf emerged as a novel virtual conference in response to the COVID-19 pandemic, aiming to provide a space for connection, idea exchange, and learning amidst widespread event cancellations. Organized by Gremlin, the event quickly gained momentum with 8,680 sign-ups, exceeding initial expectations by over four times. Despite technical challenges, the conference successfully featured thirteen speakers, drawing an engaged audience with an average participation of 3.5 hours. The event prioritized empathy, community collaboration, and quality content, allowing speakers to share their insights while fostering a sense of community through interactive platforms like Slack. Though post-event feedback highlighted concerns about email volume, the conference was largely well-received, with 97% of attendees stating it met or exceeded their expectations. The event exemplified resilience and adaptability in a rapidly changing environment, providing a blueprint for future virtual gatherings.
May 05, 2020 1,568 words in the original blog post.