November 2018 Summaries
7 posts from PagerDuty
Filter
Month:
Year:
Post Summaries
Back to Blog
DevSecOps has significantly influenced engineering culture by promoting collaboration between development, operations, and security teams, exemplified by PagerDuty's approach of integrating security into the development lifecycle. While traditional security teams have been resistant to rapid changes due to the risks involved, DevSecOps encourages a shared focus on customer outcomes and security accountability by fostering a unified language for evaluating risk versus value. PagerDuty, in partnership with security solutions like Sumo Logic, Dome9, Threat Stack, and Twistlock, aims to enhance this collaborative culture by providing real-time monitoring and actionable alerts to streamline incident response and compliance management. These integrations allow organizations to address security issues proactively without hindering the pace of development, ensuring that security enables rather than obstructs customer service. Through partnerships and strategic integrations, PagerDuty advocates for a DevSecOps culture that balances the need for speed with the imperative of security, inviting collaboration and shared responsibility across all teams involved.
Nov 28, 2018
961 words in the original blog post.
PagerDuty, founded by former Amazon employees, has long supported AWS users by transforming signals into actionable insights. The company has announced new integrations with AWS services like CloudWatch Events, GuardDuty, CloudTrail, and Personal Health Dashboard at AWS re:Invent in Las Vegas. These integrations enhance existing capabilities, allowing teams to automate operations and improve security through seamless communication with systems like Jira and ServiceNow. By leveraging CloudWatch Events, PagerDuty users can monitor AWS environment changes in real-time, while GuardDuty integration focuses on security by automating response workflows. The CloudTrail integration aids in compliance and governance by automating actions based on event histories, and the Personal Health Dashboard integration offers personalized insights into AWS service performance and availability. These developments enable teams to maintain operational continuity, optimize resource utilization, and address potential security threats efficiently.
Nov 26, 2018
1,463 words in the original blog post.
In the demanding environment of IT operations, personnel face constant pressure to minimize business disruptions, which can lead to burnout and negatively impact customer experience. A study of 85,000 services showed that those integrated with AWS experienced fewer notifications, particularly during off-hours, suggesting greater operational efficiency and reduced alert fatigue, though the exact reasons remain speculative. To improve operations health, best practices include conducting analyses of transient notifications to reduce false alarms, implementing alert grouping for better situational awareness, and maintaining consistent service taxonomies to expedite incident response. PagerDuty's Operations Health Management Service (OHMS) provides solutions by analyzing organizational health through human factors, offering actionable recommendations to enhance operational health continuously and measurably.
Nov 26, 2018
836 words in the original blog post.
Black Friday, marking the start of the holiday shopping season in the United States, presents significant challenges for retail and IT workers due to the surge in consumer activity both in stores and online. This period is characterized by an increase in system demands, leading to potential service outages that can impact customer loyalty if not managed effectively. With a substantial portion of consumers now shopping via mobile devices, ensuring the resilience and performance of digital operations is crucial. Retailers must employ strategies such as automation, effective incident management, and alert consolidation to maintain system uptime and meet customer expectations. These measures not only enhance customer satisfaction by providing a consistent shopping experience but also allow IT teams to enjoy work-life balance during the hectic holiday season.
Nov 21, 2018
944 words in the original blog post.
In 2017, PagerDuty open-sourced its Incident Response Documentation to help others avoid the initial challenges they faced, drawing on a framework used by first responders. While the original documentation was comprehensive, it was criticized for lacking detailed training methods. To address this, PagerDuty has now updated the documentation to include their Incident Response Training course, which introduces the role of the Incident Commander and is based on the training their employees receive. The course materials, including slides, notes, and a downloadable PDF, are available on the open-source site, and a recording from a 2017 event is also provided to illustrate the presentation style. This enhancement aims to serve as a practical resource for organizations looking to improve their incident response capabilities, and the updated content remains in the same GitHub repository, inviting community contributions. For those interested in more extensive training, PagerDuty University offers various programs, including private and public courses, which can be inquired about via email.
Nov 13, 2018
418 words in the original blog post.
Dealing with software bugs can be a challenging task, but Sentry, an open-source error tracking platform, offers best practices to make error resolution more manageable, including an integration with PagerDuty for efficient incident response. Key strategies include identifying the issue's impact using Sentry's tagging system, retracing users' steps with automatic breadcrumbs, and gaining context from stack traces enhanced with un-minified source code and stack locals. Involving the original developer when necessary and prioritizing alerts to avoid unnecessary notifications are also emphasized. Sentry's integration with source code management platforms helps identify suspect commits and suggest the appropriate developer for bug fixes. With tools like alert rules and event management, users can customize notifications, ensuring that only significant issues trigger alerts. These practices, combined with Sentry's and PagerDuty's capabilities, provide a comprehensive approach to error tracking and resolution, empowering development and operations teams to have a complete view of errors and manage notifications effectively.
Nov 08, 2018
820 words in the original blog post.
In an era where digital technologies are integral to business operations, disruptions in IT services can significantly impact customer satisfaction and business continuity. The Real-Time Operations Maturity Model developed by PagerDuty is designed to help organizations assess and improve their ability to handle real-time work efficiently, categorizing them into four levels: Reactive, Responsive, Proactive, and Preventative. The model emphasizes the importance of learning from past incidents, automating processes, and empowering employees with the knowledge and authority to prevent and resolve issues. A survey conducted by IDG highlights that while many organizations have progress to make in achieving real-time operations maturity, those that do so experience reduced incident rates, quicker resolutions, and lower employee attrition due to less burnout. Mature organizations benefit from faster incident acknowledgment and resolution, less downtime, and ultimately, a positive impact on their business and customer satisfaction.
Nov 07, 2018
1,247 words in the original blog post.