Home / Companies / Gremlin / Blog / August 2023

August 2023 Summaries

4 posts from Gremlin

Filter
Month: Year:
Post Summaries Back to Blog
Reliability programs are crucial for organizations to proactively manage and improve system resiliency and availability, and they should be built around four key pillars: leadership and strategy, clear ownership and handoffs, measurement and metrics, and processes and policies. These programs require more than just technology; they necessitate organizational coordination and clear strategies, goals, and accountability. Leadership buy-in, clearly defined responsibilities, and the ability to measure progress against business-relevant metrics are essential for success. Establishing consistent and robust processes and policies helps ensure ongoing compliance and improvement. Gremlin, a company specializing in reliability, advocates for these principles and offers tools and resources to help organizations uncover and address reliability risks before they impact users, including a free trial of their platform to identify hidden system risks.
Aug 31, 2023 1,541 words in the original blog post.
Gremlin has introduced a new feature called Detected Risks, designed to help organizations identify and fix common infrastructure vulnerabilities without the need for extensive testing or Chaos Engineering experiments. Detected Risks aims to enhance system reliability by providing immediate insights into potential failure points and guiding teams through remedies with minimal configuration. This tool targets prevalent issues such as zone redundancy and autoscaling misconfigurations, which are often easy to fix but can significantly impact system uptime if left unaddressed. Gremlin's approach allows engineering teams to prioritize and address these risks efficiently, demonstrating progress and improving reliability. Detected Risks offers an expanding library of risk scenarios, starting with core Kubernetes risks, and plans to continue evolving to meet broader reliability needs. The initiative underscores Gremlin's commitment to empowering businesses to build more reliable software and modernize their reliability practices at an enterprise scale.
Aug 30, 2023 1,123 words in the original blog post.
Gremlin has launched the Enterprise Chaos Engineering Certification (GECEC) program, building on the success of its previous Chaos Engineering certifications. The GECEC is designed to address the evolving needs of enterprise organizations by consolidating elements from the GCCEP and GCCEPro certifications into a more streamlined, comprehensive course that covers essential concepts like fault injection and conducting safe experiments on production systems. This new certification reflects the growing importance of Chaos Engineering in enhancing system reliability and adapts to industry changes by emphasizing automation, integration, and scalability, without requiring prior Gremlin or Chaos Engineering experience. Participants can access a variety of learning materials and receive a certificate upon completion, which can be shared on professional networks. The program targets professionals, especially Site Reliability Engineers, who wish to demonstrate their expertise in reliability and resilience in modern applications.
Aug 23, 2023 914 words in the original blog post.
Gremlin utilizes its own platform to enhance software reliability by conducting Chaos Engineering experiments that identify and address potential reliability risks. This involves a structured approach with five best practices, which include fine-tuning monitoring systems, integrating reliability tests throughout development stages, starting with tests for common failure modes, scheduling regular tests to minimize disruptions, and maintaining regular meetings to ensure issues are addressed promptly. By employing these strategies, Gremlin aims to detect and resolve system vulnerabilities before they impact users, thereby improving system stability and reducing the frequency of incidents. The platform's pre-built reliability tests and scoring system aid in systematically defining and measuring progress toward reliability standards across organizations.
Aug 07, 2023 1,903 words in the original blog post.