August 2024 Summaries
2 posts from Gremlin
Filter
Month:
Year:
Post Summaries
Back to Blog
Resilient and reliable IT systems are crucial for modern businesses, particularly in the financial sector, where outages can have significant consequences. Recent global regulations such as the Digital Operational Resilience Act (DORA) in the EU, APRA CPS 230 in Australia, and FCA PS21/3 in the UK have been introduced to enhance IT operational resilience and risk management. These regulations require companies to perform regular operational resilience testing, document and test Business Continuity and Disaster Recovery plans, verify monitoring and incident response processes, map third-party dependencies, and meet performance obligations. Gremlin offers a platform that helps organizations automate and standardize these processes by simulating failures in a controlled environment, providing documentation and reporting for compliance. This enables companies to transition from manual checklists to automated, verifiable compliance, ensuring their systems maintain minimum operational standards during failures, thereby supporting both compliance and operational resilience.
Aug 29, 2024
2,149 words in the original blog post.
Cloud dependencies, particularly managed services provided by companies like AWS, pose both opportunities and challenges due to their external control and potential for failure, which can lead to cascading effects on applications relying on them. To mitigate these risks, businesses can utilize tools like Gremlin to simulate service outages and assess their systems' resilience through the Well-Architected Cloud Test Suite, which evaluates handling of slow or unresponsive connections and certificate issues. By leveraging techniques such as caching, asynchronous calls, and redundant services, organizations can enhance reliability and prepare for eventualities such as service failures. Additionally, Gremlin's platform allows for proactive testing and monitoring of services to maintain high availability and performance, ensuring that systems are resilient to disruptions in cloud-based dependencies.
Aug 01, 2024
2,088 words in the original blog post.