Home / Companies / Gremlin / Blog / July 2026

July 2026 Summaries

2 posts from Gremlin

Filter
Month: Year:
Post Summaries Back to Blog
The Gremlin app for Dynatrace integrates resilience testing and reliability scoring into the existing observability framework of Dynatrace, offering a comprehensive view of system performance and potential failure scenarios. This integration allows engineering teams to conduct Gremlin reliability tests within the Dynatrace environment, leveraging its metrics and alerts to monitor and score system reliability in real time without additional setup. By combining Dynatrace's real-time visibility into system metrics with Gremlin's forward-looking reliability scores, organizations can anticipate risks and validate system resilience, thus enhancing their ability to manage and report on reliability investments effectively. This integration supports a seamless workflow for users by presenting reliability scores and test results directly on Dynatrace dashboards, helping executives and stakeholders make informed strategic decisions about reliability efforts. The app is available in the Dynatrace Hub, requiring a Gremlin account and specific configurations to start measuring and improving system reliability across various infrastructures.
Jul 28, 2026 1,254 words in the original blog post.
In the quest for cloud resilience and uptime reliability, managing hundreds of applications across platforms like AWS, Azure, and GCP is challenging due to the complexity of identifying failure points manually. Gremlin's automated reliability platform addresses this by expanding its Detected Risks feature, which now flags high-priority reliability risks in Azure and GCP environments, complementing its existing capabilities for Kubernetes and AWS. This feature identifies potential blind spots such as misconfigured deployments, lack of redundancy in availability zones, and missing readiness probes, which can lead to outages and compromised service availability. By automatically detecting these risks, Gremlin enables proactive risk management, allowing organizations to eliminate issues before they affect users, ensuring optimal configuration and infrastructure health. The platform also highlights the importance of features like topology spread constraints, PodDisruptionBudgets, autoscaling, and SSL certificate management to enhance service reliability. Gremlin's approach shifts reliability management from reactive troubleshooting to proactive detection, helping organizations maintain high availability and performance across cloud environments.
Jul 14, 2026 1,830 words in the original blog post.