Home / Companies / Gremlin / Blog / December 2024

December 2024 Summaries

4 posts from Gremlin

Filter
Month: Year:
Post Summaries Back to Blog
In 2024, Gremlin focused on enhancing its reliability testing platform with a series of new features and updates, including the introduction of two new experiments—Process Exhaustion and GPU stress tests—to help organizations test system resilience against concurrent workloads and GPU-based operations. The platform also unveiled a streamlined onboarding process for AWS users, allowing automatic service discovery and monitoring integration with CloudWatch metrics. Additionally, Gremlin introduced Intelligent Health Checks and AWS-specific Detected Risks to enhance AWS workflow reliability. Other improvements included enhanced agent capabilities for large Kubernetes clusters and systems with over 64 CPUs, as well as new support features for serverless and containerized workloads. The platform also incorporated customizable role-based access controls, improved dependency detection, and enhanced experiment behaviors, ensuring smoother deployment and management of reliability tests across diverse environments. Gremlin's updates reflect its commitment to helping organizations proactively identify and mitigate reliability risks, with a focus on integrating seamlessly with AWS services and enhancing user experience through UI improvements and comprehensive auditing tools.
Dec 18, 2024 2,879 words in the original blog post.
The Gremlin Service Mesh Extension, currently in private beta, enhances resilience testing by allowing more granular application-level targeting within service meshes like Istio, which are widely used in dynamic, service-based architectures such as Kubernetes. Traditional infrastructure-based testing targets specific IP addresses, but this approach lacks precision in service mesh environments where IPs frequently change due to their inherent resilience capabilities. The new extension uses the HTTP path for targeting, enabling specific interactions within an application to be tested rather than all traffic from a service, providing a more detailed understanding of application reliability. The beta includes experiments such as Network Latency, Blackhole, and Unexpected Service Response Codes, which can be used to simulate various failure scenarios and validate recovery systems. Compatible with Istio 1.22.x, the extension can be configured via the Gremlin UI or API and promises to enhance the reliability and resilience of service mesh applications without adding extra points of failure.
Dec 04, 2024 755 words in the original blog post.
As the year 2024 concludes, Gremlin has introduced several new features aimed at enhancing reliability in serverless and AI environments. The company now supports the Istio service mesh, allowing developers to run experiments on service mesh applications to simulate network issues and failures. A new GPU experiment helps test GPU-based workloads to identify potential failures before they affect users, an essential capability given the growing investment in GPU technologies. Additionally, Gremlin offers AWS PrivateLink integration for secure connections without using the public internet, along with improved onboarding for Kubernetes clusters through Argo Rollout support and auto-generated Helm commands. Updates to Linux and Windows agents enhance performance and stability for enterprise deployments, including support for systems with more than 64 processors and earlier Linux kernels. These developments are part of Gremlin's ongoing effort to empower users to detect and address availability risks proactively.
Dec 04, 2024 993 words in the original blog post.
Gremlin's GPU experiment is designed to test the reliability and performance of AI models by simulating intensive GPU workloads, thereby identifying potential failures and optimizing resource management. The experiment stresses the GPU's processing unit to its limits using OpenCL, allowing organizations to validate their systems' scalability, capacity planning, and fault tolerance. This is particularly relevant given AI's growing reliance on GPUs for parallel processing, as seen with large language models like ChatGPT and DALL-E. The experiment can simulate various scenarios, such as heavy loads, insufficient memory, or noisy neighbor impacts, and provides insights into infrastructure resilience by combining with other Gremlin experiments like blackhole scenarios. Gremlin's platform enables companies to proactively address availability risks, offering a 30-day free trial for new users to explore these capabilities.
Dec 02, 2024 1,511 words in the original blog post.