Home / Companies / PagerDuty / Blog / April 2014

April 2014 Summaries

11 posts from PagerDuty

Filter
Month: Year:
Post Summaries Back to Blog
On April 14th, PagerDuty experienced a 30-minute outage affecting both mobile and web applications, resulting in delayed alerts and account management issues for customers. The incident was caused by an increased workload on their event processing system, which led to performance degradation and timeouts in an upstream system with a retry policy, ultimately causing significant system load and availability issues. Despite the delays, no events were lost, and all alerts were eventually sent. In response, PagerDuty's operations and engineering teams quickly alleviated the problem by removing duplicate queued events and adjusting the retry policy to prevent future occurrences. Long-term solutions include rebalancing timeout and retry policies and separating event processing from customer-facing applications to enhance reliability and performance. The company has apologized for the service disruption and is committed to preventing similar issues in the future.
Apr 28, 2014 352 words in the original blog post.
PagerDuty emphasizes reliability in delivering alerts by conducting End-to-End SMS Provider Testing to proactively identify and address potential delays or outages from third-party carriers, even when their status pages indicate full availability. This initiative involves continuously sending test SMS messages through their various providers, using a system of Android phones with different mobile carrier networks, to evaluate delivery times and ensure optimal performance. If a provider is deemed degraded—characterized by delivery latencies over three minutes or multiple missed messages—the team is alerted to replace them, thereby maintaining the integrity of customer alerts. While the current process of adjusting provider priority is manual, PagerDuty plans to automate it with a probabilistic model to minimize the noise of failure alerts and focus on solving issues. This rigorous testing and automation have not only enhanced the reliability of their services but also provided deep insights into the performance of connected systems, ensuring an improved user experience.
Apr 28, 2014 684 words in the original blog post.
PagerDuty offers comprehensive integration with over 75 out-of-the-box tools, including monitoring and business applications such as chat programs and customer support portals, enhancing system visibility and incident management. Recently, PagerDuty has expanded its capabilities to include integration with Single Sign-On (SSO) providers, offering enterprise customers streamlined access management. By partnering with SSO providers like Okta, OneLogin, and Ping Identity, PagerDuty allows organizations to manage user credentials centrally, simplify login processes, and ensure secure and efficient user provisioning and deprovisioning. This integration enables a unified access point for all applications, reducing the need to remember multiple passwords and providing a single-click login experience across devices. Such collaborations help companies like LinkedIn, Netflix, and McDonald's maintain secure access to their critical applications while also managing access efficiently through a centralized identity management system.
Apr 24, 2014 683 words in the original blog post.
Website monitoring involves testing and verifying that end-users can access and use online services effectively, with the aim of identifying issues before they cause major disruptions. While simple pings can alert teams to downtime, a more comprehensive monitoring approach includes external checks, such as pinging a site every 15 seconds, and internal checks that examine the entire system stack to identify root causes of outages. Tools like PagerDuty provide a robust solution by not only tracking uptime but also monitoring event flow, processing times, and alert volumes to ensure systems function smoothly. Effective monitoring goes beyond basic checks to ensure correct content delivery, including CSS and scripts, and can trigger alerts of varying severity depending on the issue detected. Employing multiple monitoring tools can add redundancy, ensuring no alert goes unnoticed, and PagerDuty offers various integrations to enhance monitoring capabilities.
Apr 22, 2014 585 words in the original blog post.
PagerDuty's recent launch of the Multi-User Alerting feature has seen high adoption and positive feedback, addressing a major customer request by allowing incidents to be assigned and acknowledged by multiple users. This enhancement required significant architectural changes to ensure no service downtime, involving collaboration among various teams and a carefully planned incremental rollout. The new feature enhances incident management by tracking multi-acknowledgments and implementing acknowledgment timeouts, preventing incidents from being forgotten and avoiding unnecessary notifications. The alerting pipeline was adjusted to ensure "minute zero" notifications, guaranteeing alerts are sent even if incidents are quickly resolved or acknowledged. Multi-User Alerting eliminates the need for previous workarounds, enabling customers to configure escalation policies more efficiently by adding up to 10 escalation targets, ensuring timely notifications reach the right personnel.
Apr 16, 2014 743 words in the original blog post.
Continuous integration (CI) is a practice in software development where team members frequently merge their work to reduce conflicts and enhance the quality and reliability of software. At PagerDuty, this involves automated builds and tests with a focus on detecting and fixing bugs quickly, supported by a test-driven development approach. The process begins with creating JIRA tickets for collaboration and tracking, followed by branching from a distributed version control system like Git, which enhances redundancy and local development. Tests are prioritized for security, strategic changes, consistency, and shared knowledge, and they are classified into semantic, unit, functional, integration, and load tests. These tests ensure code quality and reliability before deployment. The deployment process includes manual peer reviews and semi-automatic deployment using tools like Capistrano and Chef, accompanied by notifications to keep the team informed and avoid concurrency issues. Overall, CI helps maintain a baseline of software quality, reducing risks associated with releases.
Apr 14, 2014 1,402 words in the original blog post.
On March 25th, PagerDuty experienced a significant service degradation lasting three hours, affecting customers by delaying 11% of notifications and failing to accept 2.5% of event attempts. The issue stemmed from an overload in their Cassandra-based notifications pipeline, exacerbated by both steady-state and bursty workloads from scheduled jobs. Although the system's retry logic was designed to handle transient failures, it inadvertently prolonged the overload period. To address these issues, PagerDuty plans to temporally distribute and flatten scheduled job loads, isolate systems onto separate Cassandra clusters to prevent cross-system interference, and adjust failure detection and retry policies to better handle overloads. They are committed to enhancing reliability and will incorporate overload scenarios into their failure testing regime.
Apr 11, 2014 590 words in the original blog post.
PagerDuty utilizes webhooks to enable seamless communication between applications, offering a streamlined alternative to complex APIs for integrating with platforms like HipChat, Slack, and Zapier. Webhooks facilitate custom integrations, as demonstrated by Dave Hayes' hackday project which used Firebase for animated incident mapping. In another project, webscript.io was employed for its simplicity and versatility, allowing for the automation of incident conference calls through webhooks that trigger scripts, effectively enhancing team collaboration during high-severity incidents. The project initially considered using Twilio but opted for VoiceChatAPI by Plivo for its ease of creating conference calls without needing an API key. This innovative approach highlights webscript.io's utility for various functions, such as monitoring device status, checking website uptime, and converting webhooks into API calls or email alerts. The project was supported by contributors from Webscript.io and Plivo, with additional resources available in the pagerduty-webscripts GitHub repository for users interested in further customization and integration opportunities.
Apr 10, 2014 543 words in the original blog post.
PagerDuty has enhanced its escalation policies by introducing Multi-User Alerting, allowing up to 10 team members to be notified simultaneously at each level of an escalation policy. This feature ensures rapid response to high-severity incidents by alerting multiple responders, including primary, secondary, and tertiary, thereby minimizing the chances of oversight. The system also facilitates the integration of new team members through shadowing, where trainees can share on-call duties with experienced engineers. Non-responders, such as managers or interested parties, can stay informed in real-time by adding themselves to escalation policies, and multiple teams can now be notified for incidents requiring cross-functional collaboration. The feature, already available, aims to streamline incident management without unnecessary escalations, encouraging responsible alerting practices.
Apr 08, 2014 417 words in the original blog post.
PagerDuty has proven to be an essential tool for various companies transitioning to or operating within a DevOps model by streamlining alert management and on-call scheduling, leading to improved efficiency and team collaboration. Brightcove adopted PagerDuty to manage on-call schedules effectively and foster a sense of code ownership among developers, while Sumo Logic uses it to quickly onboard new hires unfamiliar with DevOps environments. MLS Digital relies on PagerDuty to prevent burnout by ensuring fair distribution of on-call responsibilities and enabling team members to maintain a work-life balance. At Ping Identity, PagerDuty has become integral to their infrastructure, facilitating cross-functional collaboration and significantly reducing incident repair times. Simple employs PagerDuty across different non-engineering teams, such as PR, to maintain communication and respond promptly to outages, enhancing their operational workflow.
Apr 07, 2014 468 words in the original blog post.
Transitioning to a DevOps culture involves embracing a collaborative environment where developers are empowered through self-service tools, which enable them to handle issues independently, thereby enhancing productivity and maintaining consistency across various environments. In a DevOps model, tools like Chef are utilized to treat infrastructure as code, incorporating version control, automated testing, and peer reviews to ensure reliability. Prioritizing monitoring metrics based on customer importance ensures that alert systems are responsive, minimizing service disruptions and improving customer satisfaction. Additionally, fostering connections among team members and between people and systems is crucial, with platforms like GitHub facilitating collaboration, knowledge sharing, and accountability. Implementing an on-call strategy helps align developer responsibilities with customer needs, although initially met with resistance, it ultimately enhances team cohesion. Achieving a successful DevOps transformation relies on breaking down departmental silos and focusing on shared goals, self-service empowerment, and effective communication.
Apr 03, 2014 760 words in the original blog post.