May 2017 Summaries
15 posts from PagerDuty
Filter
Month:
Year:
Post Summaries
Back to Blog
In the complex realm of IT, while monitoring tools are essential for gathering data on applications and systems, the challenge lies in transforming this information into actionable intelligence and embedding effective response processes within the IT organization. Despite the widespread deployment of monitoring systems, many alerts are minor and often ignored, leaving organizations unprepared for handling significant alarms that indicate potential catastrophic failures. The integration of modern monitoring tools with IT incident resolution platforms, through APIs, facilitates the triangulation of alarms, helping to identify root causes and analyze data centrally to prevent recurring issues. This integration is crucial in reducing the cognitive load on IT teams and correlating application performance with business metrics like revenue loss and customer churn. However, a significant portion of IT professionals find increased complexity makes their jobs more challenging, with many not fully monitoring their networks. It is essential to move beyond merely gathering affected parties in a "war room" to blame one another and instead implement structured processes and best practices for incident resolution. These practices ensure rapid issue resolution, often without convening meetings, by using embedded runbooks and automated troubleshooting, thereby maximizing the value of IT monitoring and minimizing downtime and blame.
May 31, 2017
695 words in the original blog post.
Zayna Shahzad, a Software Engineer at PagerDuty, shares her experience shadowing the Customer Support team to foster empathy and understand different departmental roles within the company. During her day of shadowing, she gained insights into the systematic processes of the support team, their approach to ticket management, and the importance of collaboration and communication. Shahzad observed the team's dedication to customer satisfaction, their commitment to continuous learning, and their ability to manage complex queries efficiently. Her experience highlighted the value of stepping outside one's immediate team to appreciate the diverse contributions across the company, reinforcing PagerDuty's culture of empathy and teamwork. The initiative to have employees from other departments, including Engineering and Product, participate in similar shadowing experiences has proven beneficial in enhancing empathy and understanding of customer needs.
May 25, 2017
1,650 words in the original blog post.
PagerDuty's engineering teams are committed to Agile development principles, prioritizing rapid iteration and direct communication over extensive written specifications, while recognizing the uniqueness of individual teams by supporting a flexible approach to Agile practices. Initially, new teams often adopt the Scrum framework, engaging in daily standups, sprint planning, and Storytime meetings to foster teamwork and set shared goals. However, PagerDuty encourages continuous improvement, allowing teams to adapt and modify Scrum practices to suit their specific needs, which may involve altering workflows, adjusting meeting schedules, or even transitioning to different methodologies like Kanban. The company emphasizes empowerment, enabling teams to experiment with and refine their processes to enhance productivity and satisfaction, rather than adhering strictly to standardized guidelines. Ultimately, PagerDuty values adapting tools and processes to empower its teams, ensuring they remain effective and content in their work environments.
May 23, 2017
667 words in the original blog post.
PagerDuty's SVP of Product Development, Tim Armandpour, emphasizes the importance of adopting robust incident response processes to ensure system reliability in a world that demands constant availability. He introduces the concept of "Failure Fridays," where their engineering team deliberately injects failures into their live production environment to improve system resilience and practice effective incident response. This approach focuses on understanding failure scenarios, fostering collaboration across organizational parts, and preparing teams to handle real-life incidents without panic. Key learnings include testing various failure scenarios to expose vulnerabilities, maintaining a blameless post-mortem process to derive actionable improvements, and treating each identified vulnerability as a chance to enhance infrastructure resilience. The practice underscores that reliability is crucial to their customer promise, and preparing for failures is a vital part of their operational strategy.
May 22, 2017
452 words in the original blog post.
PagerDuty's blog post delves into the evolving role of security within the DevOps framework and the shifting dynamics between development and operations teams. During a webinar featuring experts Ilan Rabinovich, Chris Gervais, John Rakowski, and Arup Chakrabarti, the discussion centered on the integration of security as an inherent component of the DevOps model, emphasizing its importance as part of good software hygiene and the necessity of involving security teams early in the development process. The panelists argued that security should transition from being a separate entity to a shared responsibility within DevOps, fostering collaboration and ensuring it becomes the "domain of the many." Additionally, the conversation explored the idea of central Ops teams moving closer to the application codebase, highlighting the benefits of shared knowledge and closer interaction between Ops and Dev to enhance efficiency and problem-solving capabilities. The experts underscored the importance of breaking down silos and adopting automation to streamline operations, ultimately leading to faster issue resolution and improved business outcomes. As the DevOps landscape continues to evolve, the integration of security and closer collaboration between Dev and Ops are seen as essential to driving innovation and maintaining agility.
May 18, 2017
1,927 words in the original blog post.
PagerDuty is hosting the Digital Operations Excellence awards at the PagerDuty Summit 2017 to honor outstanding achievements in digital operations management by its customers and partners. The awards recognize innovative and transformative approaches, with categories such as Scaling Digital Operations Management, Innovation, Transformation, Best Security Incident Response, Best Customer Support Use Case, Customer Experience, and Kick Start for startups. Partner awards include Innovation and Transformation categories, highlighting excellence in delivering value and leading digital transformation. Finalists were notified on August 1, 2017, with winners announced on stage at the summit in San Francisco on September 7, 2017.
May 17, 2017
625 words in the original blog post.
A diverse group of industry leaders shared their experiences and plans for digital transformation, highlighting challenges such as transitioning to the public cloud, modernizing central operations, and refactoring code from monoliths to microservices. These leaders, representing sectors like technology, media, and finance, are focused on improving agility to enhance innovation in customer-facing services by adopting DevOps practices and a developer ownership model. Despite the varied contexts, all emphasize the importance of a supportive culture of collaboration and continuous learning in driving successful transformation. The narrative underscores that while tools are essential, they are insufficient without a cultural shift, echoing the sentiment that transformation is as much about changing "hearts and minds" as it is about technological advancement. PagerDuty aims to facilitate these efforts by providing products that support best practices and enable teams to operate autonomously with consistent processes, particularly in incident management and DevOps practices.
May 16, 2017
553 words in the original blog post.
The new PagerDuty and HipChat extension enhances collaboration by allowing responders to manage incidents directly from their chat window, offering improved speed and productivity for modern development and operations teams. The updated v2 HipChat extension, built from scratch, simplifies setup by eliminating the need for copying integration keys or switching between browser tabs, thus streamlining the integration process. Once configured, users receive color-coded incident updates—red for triggered, yellow for acknowledged, and green for resolved—directly in their HipChat rooms, enabling real-time monitoring and action without needing to switch to the PagerDuty web app. Responders can acknowledge or resolve incidents through the HipChat sidebar and use slash commands to update multiple incidents simultaneously, with further integrations expected to improve functionality without needing additional key setups.
May 15, 2017
418 words in the original blog post.
PagerDuty has launched the "Succeed at PagerDuty" webinar series to provide in-depth product insights and live Q&A sessions, aiming to address frequently asked customer questions and enhance platform utilization. These 30-minute webinars, available live and on-demand, will occur approximately once a month and cover topics highly requested by users. The series includes sessions on approaching service groups, on-call scheduling, and using team-based permissions, each led by a PagerDuty Customer Success Manager and designed to offer practical best practices and solutions. The initiative underscores PagerDuty's commitment to customer support and success by facilitating better understanding and usage of its platform features.
May 11, 2017
446 words in the original blog post.
Incident lifecycle management is a proactive framework designed to efficiently handle and resolve incidents in software and IT companies, minimizing service disruption and stress for incident-response teams. Rooted in the ITIL model, which emphasizes maintaining customer services, this framework involves several key phases: the initial response where alerts are logged and categorized; Level 1 response teams that address issues with known solutions and maintain communication with affected clients; and Level 2 teams that handle more complex problems and may involve third-party support. Post-resolution processes include verifying and documenting the incident's resolution and learning from it to prevent future occurrences. Additionally, the management of major incidents and the use of temporary workarounds are crucial elements, as they prioritize customer service restoration but also highlight the importance of replacing quick fixes with long-term solutions to avoid accumulating technical debt. By implementing a tailored incident lifecycle management framework, organizations can ensure reliable service continuity, reduce chaos, and enhance their long-term success.
May 10, 2017
1,045 words in the original blog post.
PagerDuty has introduced integrated support for postmortems within its incident management platform, aiming to enhance the cycle of learning and improvement following major incidents. This new feature streamlines the postmortem process by simplifying the construction of incident timelines and integrating relevant PagerDuty and chat activity, allowing for efficient analysis of root causes and response effectiveness. The postmortem editor guides users in summarizing incidents, identifying underlying issues, and determining actionable improvements. Additionally, the platform offers a catalog for easy access to postmortems and customizable templates, facilitating a culture of shared learning and continuous improvement. This functionality is available to customers on Standard and Enterprise plans, with resources such as a free postmortem handbook offered to further support effective postmortem practices.
May 09, 2017
745 words in the original blog post.
PagerDuty has introduced new functionalities aimed at enhancing the Incident Resolution Lifecycle by helping organizations differentiate between major incidents and routine operational issues, thereby streamlining incident resolution and learning processes. The updates include stages such as Assess, Respond, and Learn, which allow responders to quickly identify incident impacts, coordinate responses across teams, and efficiently build postmortem timelines. As digital operations grow increasingly complex, the need for consistent practices and roles in incident response becomes paramount, with PagerDuty facilitating this through a framework that distinguishes major incidents and integrates with tools like ServiceNow and JIRA to eliminate redundant efforts. The new features, including an incident postmortem builder and expanded integration capabilities, aim to foster a culture of continuous learning and improvement, making it easier for organizations to manage and scale their incident resolution processes. These advancements are part of PagerDuty’s broader effort to support operational maturity by automating and simplifying tools to enhance the overall customer experience during outages.
May 08, 2017
854 words in the original blog post.
Incident management should extend beyond simply resolving infrastructure issues by leveraging historical data to proactively prevent future incidents and enhance system resilience. By standardizing and centralizing incident data from various monitoring systems, organizations can overcome challenges such as varied data formats and limited historical records. Tools like Logstash, Splunk, Papertrail, and PagerDuty facilitate the collection and standardization of data, enabling the identification of patterns and trends through visualizations. Effective data analysis involves understanding metrics like incident frequency, mean time to acknowledge and resolve, team workload distribution, and alert generation by monitoring systems. By addressing these elements, organizations can avoid repeating past mistakes, optimize their infrastructure, and transform incident management from a reactive to a preventive approach, embodying the principle that prevention is more effective than cure.
May 04, 2017
1,064 words in the original blog post.
PagerDuty and Atlassian are enhancing their partnership to streamline incident resolution processes by integrating their tools, as seen with the successful launch of the PagerDuty HipChat Extension, which allows users to manage incidents directly within HipChat. Building on this positive reception, they announced a new integration between PagerDuty and JIRA Software Cloud aimed at providing a seamless experience for developers who are increasingly responsible for their code in production. This integration includes features like easy setup, one-click JIRA ticket creation from PagerDuty incidents, and linked navigation between the systems to improve efficiency and reduce resolution times. The collaboration underscores a commitment to delivering innovative solutions and improving customer experiences by aligning their products more closely, with ongoing efforts to support the developer community and enhance the incident resolution lifecycle.
May 02, 2017
404 words in the original blog post.
PagerDuty's platform allows users to creatively extend its functionalities, demonstrated by a customer's suggestion to integrate video conferencing into their incident response process. By embedding a video conference tool, like Appear.in, into the Operations Command Console, users can enhance their coordination during incidents, although the solution shared is a simple overview rather than the exact implementation used. This flexibility showcases the platform's adaptability, as it supports the integration of various tools, such as chats and status pages, through its custom URL module feature in beta testing. The platform encourages users to explore and share unique ways of enhancing their incident response capabilities, fostering a community of innovation and collaboration.
May 01, 2017
277 words in the original blog post.