January 2022 Summaries
13 posts from PagerDuty
Filter
Month:
Year:
Post Summaries
Back to Blog
PagerDuty has announced a series of updates and enhancements to its platform, focusing on improving On-Call Management, Event Intelligence, Event Orchestration, and mobile products. These updates aim to help users resolve incidents more efficiently by reducing noise and manual event processing, enhancing security through recent Rundeck releases, and supporting automation. New features include the general availability of Round Robin Scheduling for distributing on-call responsibilities, an updated mobile app interface for incident response, and a decision engine for event orchestration. The platform also offers demos and webinars to showcase integrations with tools like ServiceNow and Rundeck, and invites users to explore its capabilities across various organizational departments beyond IT. Additionally, PagerDuty emphasizes community engagement through Twitch streams and encourages users to participate as design partners for future developments.
Jan 31, 2022
1,091 words in the original blog post.
In the fourth post of the EI Architecture series on Intelligent Alert Grouping, PagerDuty's Chris Bonnell discusses how service design can enhance the experience with Intelligent Alert Grouping and the PagerDuty app. The post emphasizes the importance of having a clear and actionable service definition, which should be specific enough for understanding but broad enough for organizational applicability, highlighting that services should be fully owned by a team responsible for incident response. It distinguishes between technical services, which are used for alert grouping, and business services, which are aggregates of technical services based on business logic. The balance between service granularity and ownership is explored, suggesting that services should be defined based on functionality and escalation paths. The post advises reviewing existing projects to ensure correct service grouping for effective incident management and hints at best practices available in the Full Service Ownership Ops Guide for service naming and ownership.
Jan 28, 2022
1,042 words in the original blog post.
The blog post by Chris Bonnell discusses techniques to enhance the effectiveness of Intelligent Alert Grouping (IAG) by optimizing alert titles for both human readability and machine learning processes. It emphasizes the importance of concise and clear alert titles across different notification platforms, considering character limits and readability. The post also delves into how machine learning uses natural language processing techniques, such as anonymization, tokenization, and lemmatization, to analyze alert titles for correlation and pattern recognition. By using distinctiveness and frequency to tailor alert titles, the article suggests that one can improve the IAG model's ability to accurately group alerts, while still keeping human users in mind, as they can access more detailed incident information beyond the title. The article advises a balanced approach that slightly favors machine learning optimization to ensure effective alert management.
Jan 27, 2022
1,214 words in the original blog post.
PagerDuty has introduced Event Orchestration, a new feature designed to help customers manage noise and complexity in incident response by automating repetitive tasks and reducing manual intervention. The feature builds on existing Event Rules by incorporating a decision engine within the event ingestion pipeline, allowing precise automation through complex logic and condition-based actions. This enables users to reduce noise by suppressing or routing non-critical alerts and to automate routine diagnostic steps, thus freeing up responders to focus on incidents requiring their expertise. Event Orchestration aims to enhance productivity by preemptively handling well-understood tasks before notification, with major use cases including noise reduction and automation of initial incident response phases. Users can learn more through a knowledge base, demo, and upcoming webinar featuring Frank Emery, who will discuss the feature's development and applications.
Jan 26, 2022
988 words in the original blog post.
PagerDuty has implemented comprehensive programs to support parental leave and employee wellness, reflecting its commitment to fostering a healthy workplace culture. Nearly 30% of PagerDuty employees identify as parents or caregivers, and the company’s BabyDuty program offers up to 22 weeks of paid parental leave for pregnant parents and 12 weeks for non-pregnant and adoptive parents in the US and Canada. A phased return-to-work policy and expanded financial support during leave further ease the transition back to work. The company also provides memberships to Cleo and Care.com for parental support, and backup care options are available. In response to employee stress during the COVID-19 pandemic, PagerDuty introduced Dutonian Wellness Days and a company-wide Wellness Week to promote rest and recharge, resulting in improved resilience scores among employees. Additionally, the company has enhanced its retirement matching contributions and adjusted health care costs to ease financial pressures. These initiatives demonstrate PagerDuty’s dedication to valuing its employees and continuously improving their work experience.
Jan 25, 2022
972 words in the original blog post.
PagerDuty emphasizes the importance of distinguishing between general incidents and Major Incidents, highlighting the need for robust telemetry and service relationships to effectively triage and respond to technical issues. The traditional swarming approach, which involves alerting the entire organization to an incident, is critiqued for its inefficiency, as it often results in confusion, resource wastage, and slower recovery times due to the lack of clear roles and communication. Instead, PagerDuty advocates for "Full Service Ownership," where specific teams are responsible for their services, supported by clear documentation of dependencies and escalation policies, which streamlines incident response by ensuring that knowledgeable responders are mobilized quickly. This modern approach to incident management, supported by a comprehensive service directory and strong communication plans, reduces the need for large-scale swarming, enhances efficiency, and ensures both internal and external stakeholders are kept informed, ultimately improving organizational response times and resource allocation.
Jan 18, 2022
1,779 words in the original blog post.
After two years of significant investment in cloud and related technologies to support hybrid working and digital-first models, 2022 presents a critical juncture for IT and digital leaders to maintain innovation, as research shows that companies with a focus on innovation during crises can outperform the market by 30% in subsequent years. To thrive competitively, organizations must enhance developer productivity, simplify infrastructure choices, and reduce complexity. This involves investing in new tools and cloud-based environments to boost developer efficiency, avoiding over-engineering by being pragmatic about hybrid cloud infrastructures, and managing complexity through real-time dependency management and automation. As digital services continue to expand, freeing up developers to focus on innovation rather than maintenance is essential for driving growth, with firms that excel in "developer velocity" achieving significantly higher revenue growth.
Jan 13, 2022
809 words in the original blog post.
PagerDuty highlights its commitment to customer success and its leadership in the AIOps space, as demonstrated in the Winter 2022 G2 Grid for AIOps Platforms Relationship Index. The company focuses on enhancing intelligent, automated operations by building on existing strategic capabilities without requiring organizations to overhaul their systems. Recent advancements include integrating Rundeck for automation, enhancing mobile capabilities, and introducing graphical service mapping and event orchestration to improve incident management and decision-making. PagerDuty's AIOps solution offers quick implementation, leveraging machine learning and data science to reduce noise, speed up root cause analysis, and automate processes, all while seamlessly integrating into various tech stacks with over 600 partners. The platform empowers distributed teams with self-service operations and is designed to streamline incident response through built-in automation and actionable intelligence, ensuring critical information reaches the right people promptly.
Jan 12, 2022
547 words in the original blog post.
PagerDuty has introduced Round Robin Scheduling to evenly distribute on-call shift responsibilities among team members, aiming to enhance efficiency in incident response and reduce burnout. This feature automatically assigns new incidents across different users or on-call schedules at an escalation level, preventing workload imbalance and allowing each responder to focus on individual alerts. By implementing Round Robin Scheduling, teams can set up a rotation that ensures fair distribution of incidents, thereby decreasing mean time to acknowledge (MTTA) and mean time to resolve (MTTR) as each alert receives dedicated attention. This approach also reduces the frequency of escalations by having alternative responders ready to address incoming issues, which in turn minimizes downtime and improves customer experiences. The feature is available for all Business and Digital Operations plans, and interested teams can try it for free for 14 days while current customers can upgrade to access this functionality.
Jan 11, 2022
611 words in the original blog post.
Many support organizations continue to use the traditional tiered support model, where customer issues escalate through multiple levels of a support hierarchy, typically involving three tiers. While this model is effective for handling less severe issues, it often proves inefficient for critical incidents due to delays caused by escalations and handoffs, leading to negative customer experiences. An alternative approach, Intelligent Swarming, eliminates tiered support by having the initial customer service agent maintain ownership of the ticket and collaborate with experts in real-time to resolve the issue. This method, utilized by companies like PagerDuty, emphasizes collaboration and swift response through tools and features that ensure the right experts are engaged immediately, enhancing both the customer and agent experience. As digital innovation advances, integrating such real-time, collaborative frameworks into customer service can significantly improve efficiency and satisfaction.
Jan 10, 2022
1,047 words in the original blog post.
PagerDuty’s Operations Cloud offers a versatile solution for managing critical work across various business departments beyond IT, including human resources, sales, finance, and marketing. For human resources, PagerDuty facilitates efficient offboarding and other processes by sending timely notifications, ensuring tasks are completed promptly and reducing manual follow-ups. In sales, it helps expedite the deal cycle by integrating with systems like Salesforce to ensure high-priority approvals are handled swiftly, preventing delays that could jeopardize deals. Finance teams benefit from PagerDuty's ability to send real-time alerts for failed transactions, allowing them to rectify issues quickly and avoid financial penalties. Marketing teams use it to swiftly respond to market changes and tool downtimes, optimizing campaign launches and maximizing return on investment. Overall, PagerDuty enhances workflow efficiency, reduces operational overhead, and supports seamless operations across different business functions by ensuring critical tasks are prioritized and addressed promptly.
Jan 07, 2022
1,193 words in the original blog post.
Scrum ceremonies are pivotal events within the agile methodology framework that facilitate organized collaboration and communication among team members, ultimately driving efficient project progression. These ceremonies, which include Sprint Planning, Daily Scrum, the Sprint, Sprint/Iteration Review, and Retrospective, are strategically placed at crucial points in the production process to ensure that team members maintain a clear understanding of project phases and priorities. Sprint Planning sets the stage by defining the sprint backlog and team expectations, while Daily Scrum meetings provide quick updates on task progress. The Sprint represents the actual work period where tasks are completed, followed by a Sprint/Iteration Review that showcases completed work and allows for feedback. Finally, the Retrospective focuses on refining processes to enhance future performance. By maintaining this structured approach, teams can improve production processes, deliver reliable updates, and build trust with users and stakeholders, making scrum an effective tool for nearly 60 percent of organizations engaged in product development.
Jan 06, 2022
1,035 words in the original blog post.
In a discussion involving on-call engineers from nine teams at PagerDuty, the focus was on the human aspects of being on-call, emphasizing the importance of team empathy, stress management, and suitable on-call rotations. Key takeaways include the necessity of creating a supportive team culture where it is acceptable to ask for overrides, the importance of not constantly monitoring graphs to reduce stress, and the recognition of the stress involved with postmortems, suggesting decompression time for engineers after major incidents. The discussion also highlighted that configuring low-urgency alerts can minimize overnight disturbances, and the potential burnout from week-long on-call shifts, encouraging the exploration of alternative scheduling options like shorter or split shifts. Ultimately, fostering an empathic team culture and tailoring on-call schedules to team preferences can help mitigate burnout and stress, contributing to a more sustainable on-call experience.
Jan 05, 2022
1,243 words in the original blog post.