Home / Companies / Gremlin / Blog / June 2024

June 2024 Summaries

2 posts from Gremlin

Filter
Month: Year:
Post Summaries Back to Blog
Observability and incident response are crucial for minimizing downtime and ensuring reliable software systems, but resilience testing adds a necessary layer by proactively identifying potential points of failure within complex architectures. Resilience testing works in tandem with observability to monitor system metrics and uses techniques like Fault Injection to simulate problems, allowing teams to address issues before they cause outages. It also complements incident response by verifying system resilience to known failure conditions and refining alert systems, ensuring that only critical incidents trigger responses. Integrating these practices enhances systems' reliability and availability, helping organizations meet customer demands and operational goals.
Jun 28, 2024 967 words in the original blog post.
Gremlin has launched Gremlin for AWS, a suite designed to help engineering teams improve the reliability of applications hosted on AWS by identifying and mitigating potential causes of downtime. The service introduces features like Intelligent Health Checks, which monitor service health by analyzing metrics such as throughput, latency, and error rates, and the Well-Architected Cloud Test Suite, which includes tests to ensure applications meet AWS best practices. Additionally, the platform now includes AWS-specific Detected Risks that monitor load balancer configurations to prevent incidents. Gremlin for AWS offers an automated onboarding process for services, making it easier for teams to implement resilience testing without extensive manual setup. The platform aims to make advanced reliability practices more accessible and efficient for organizations, offering a 30-day free trial for users to explore its capabilities.
Jun 20, 2024 1,275 words in the original blog post.