Home / Companies / PagerDuty / Blog / June 2017

June 2017 Summaries

15 posts from PagerDuty

Filter
Month: Year:
Post Summaries Back to Blog
Engineering and business teams often face challenges in collaboration due to their different operational methodologies, such as Agile development versus traditional project management. However, successful collaboration is possible, as demonstrated by PagerDuty during a major business system implementation involving engineering and business processes. By focusing on use case-based requirements, clearly dividing work, visualizing tasks in a shared management tool, and maintaining regular communication, teams can align their efforts to deliver compatible and valuable outcomes. This approach not only mitigates the risk of incompatible outputs but also enhances transparency and dependency management across teams, ultimately fostering a better internal system that supports business objectives. The success of this collaborative strategy at PagerDuty has led to its adoption across more teams within the organization.
Jun 30, 2017 600 words in the original blog post.
Becoming a new parent unexpectedly parallels the experience of being an on-call engineer, as highlighted by the author's journey of fatherhood, where tools and skills from his professional life at PagerDuty proved surprisingly applicable. By adopting technical strategies like setting up PagerDuty alerts for newborn care tasks, utilizing incident commander tactics to maintain calm during childbirth, and applying scribing techniques to track his baby's feeding and diaper changes, the author seamlessly integrates engineering methodologies into parenting. Despite the humorous and tongue-in-cheek narrative, the discussion delves into serious insights, such as the importance of staying calm under pressure, creating structured routines, and the challenges of functioning without the option of escalation, akin to incident response scenarios. The piece concludes by acknowledging the shared experiences and camaraderie among parents, drawing a parallel to the solidarity found in the engineering community, and underscoring the universality of the lessons learned from both roles.
Jun 29, 2017 3,583 words in the original blog post.
PagerDuty demonstrates a strong commitment to diversity and inclusion, notably by appointing women and people of color to leadership roles and fostering open discussions about workplace culture. A significant yet often overlooked aspect of PagerDuty's inclusive culture is its approach to internal technology, which facilitates effective collaboration across diverse technical backgrounds, time zones, and working styles. The company has evolved its toolchain to make local development and production deployments safe and accessible, empowering team members, including those with limited engineering experience, to contribute directly to product improvements. Additionally, PagerDuty's business analytics team has developed a robust data warehouse and analytics toolchain, making data-driven decision-making accessible to all employees and promoting a culture of empowerment. The company relies on tools such as Slack, Google Docs, and Trello for remote collaboration, which not only supports distributed work but also fosters inclusivity by accommodating various communication styles. PagerDuty also leverages its own platform to coordinate real-time responses across different business departments, emphasizing well-documented and rehearsed plans for incident management. While technology plays a crucial role, the tools are seen as products of an organizational culture that values inclusivity, as efforts are consistently made to ensure all employees can participate fully, regardless of their background or role.
Jun 28, 2017 1,493 words in the original blog post.
Attending the Girls In Tech Catalyst Conference in San Francisco proved to be a transformative experience for a summer intern at PagerDuty, offering a refreshing departure from traditional professional events with its vibrant, inclusive atmosphere. The conference featured TED Talk-style presentations by diverse female leaders from various sectors of the technology industry, including PagerDuty CEO Jennifer Tejada, who spoke on "Grace Under Pressure." Tejada emphasized that the path to leadership is fraught with sacrifices and obstacles, and she shared the importance of maintaining composure, courage, and dignity—qualities she defined as grace—in overcoming challenges. This message resonated deeply with the intern, who, as an Asian woman and Filipino citizen, reflected on personal struggles with perceived femininity and assertiveness in professional settings. The conference was an empowering testament to the resilience and determination of women who have succeeded in a male-dominated industry without compromising their aspirations. The intern left the event inspired by the strength and authenticity of these leaders and reassured about her place in the tech industry, buoyed by the supportive environment at PagerDuty and a renewed belief in her potential.
Jun 27, 2017 957 words in the original blog post.
In the era of IoT and cloud technology, managing the vast amounts of data generated by data centers poses a significant challenge similar to information overload, where too much data hampers effective decision-making. While it is tempting to collect extensive monitoring data due to low costs, this can lead to alert fatigue and inefficiencies by overwhelming staff with low-priority issues. Effective monitoring involves selectively focusing on critical events such as security incidents, host failures, and resource exhaustion, while less critical data like CPU usage and network load can be monitored without triggering alarms. Tools like PagerDuty help mitigate alert fatigue by sending notifications only to relevant personnel, and integrating log analytics tools, such as Splunk, allows for identifying trends without being overwhelmed by individual data points. This balanced approach ensures that data centers remain efficient and responsive, without succumbing to the pitfalls of excessive data collection.
Jun 22, 2017 932 words in the original blog post.
Chaos Engineering involves experimenting on distributed systems to ensure they can withstand unpredictable conditions, with companies like Netflix, Dropbox, and Twilio employing such techniques. At PagerDuty, Chaos Engineering has evolved from manual fault injection to the development of an automated fault injection system called ChaosCat, which is inspired by Chaos Monkey but is more adaptable to various service types. Initially, failures were manually injected to allow precise control and understanding, but as the infrastructure grew, automation was introduced, enabling individual teams to conduct their own fault injections. ChaosCat operates as a Scala-based Slack bot that conducts randomized chaos attacks during business hours, only when the system is fully operational, ensuring teams are ready and able to address issues as they arise. This approach has highlighted the importance of addressing gaps in run books and on-call rotations, leading to more automation and prioritization of technical debt, thereby enhancing confidence in service reliability. Although ChaosCat is currently not open-sourced due to its deep integration with PagerDuty's internal systems, the company encourages feedback and questions, hoping more organizations will adopt Chaos Engineering to test and improve the resilience of their infrastructures.
Jun 21, 2017 801 words in the original blog post.
Effective alert management is crucial for IT departments to prevent alert fatigue and maintain system efficiency by managing the overwhelming influx of notifications. This process involves filtering alerts by consolidating them into incidents and prioritizing them based on impact and urgency, often following ITIL guidelines. Impact is determined by the incident's scope on users, departments, and services, while urgency is about how quickly the problem will affect the system. By automating parts of this triage process and employing solutions like PagerDuty, organizations can suppress non-actionable alerts and focus on high-impact, high-priority incidents, minimizing distractions and effectively addressing critical issues.
Jun 20, 2017 936 words in the original blog post.
PagerDuty has announced the Public Beta launch of its Community platform, designed to foster interaction among users through forums, discussions, and shared content. The platform aims to serve as more than just a forum by encouraging collaboration and participation in various activities, such as technical Q&A, sharing best practices, and engaging with community events. The community will be further enriched by a dedicated engineering blog offering technical articles for DevOps practitioners. An AMA session with Charity Majors is scheduled to celebrate the launch, and users are encouraged to contribute their own content and join Early Access programs. The platform is integrated with PagerDuty's account system, allowing for easy sign-in, and is currently in Beta, with improvements expected before its full release. Feedback is actively sought to enhance the community experience.
Jun 15, 2017 426 words in the original blog post.
At the Mind The Product conference in San Francisco, key themes for IT Operations professionals emerged, emphasizing the need for a shift in management approaches to foster digital transformation and highlighting the importance of empathy in creating emotional connections with users. Janice Fraser from Bionic underscored the necessity of embracing cultural change and rewarding learning over certainty to drive innovation effectively, a point echoed by Zenka, while Dave Wascha from PhotoBox stressed the role of empathy in product management, urging professionals to build social connections with users. The conference also featured a light-hearted moment with the PagerDuty barbershop quartet, illustrating the value of community and humor, especially during challenging times, as a reminder that while technology issues are transient, human connections are paramount. The event, well-regarded for its contribution to the professional growth of product managers, was celebrated for its engaging and insightful sessions, with special recognition given to the organizers and participants who enriched the experience.
Jun 14, 2017 607 words in the original blog post.
PagerDuty's security team emphasizes collaboration over obstruction by developing tools and policies that facilitate secure practices while engaging with all organizational departments, as discussed in a conversation with Guy Podjarny from The Secure Developer. They strive to make security an operational concern similar to the Ops to DevOps transition, focusing on education and innovation, such as their in-house security training program that highlights password vulnerabilities and promotes password manager usage to enhance both professional and personal security. The team also discussed their experiences implementing two-factor authentication using Duo and Yubikeys, as well as leveraging operational tools like Splunk and Chef for security purposes, while acknowledging challenges in tool adoption and decision-making processes.
Jun 14, 2017 507 words in the original blog post.
PagerDuty's upcoming webinar will showcase the latest features designed to enhance incident resolution for developers and ITOps teams, emphasizing streamlined processes and automation to improve customer experiences and minimize downtime. Developers are increasingly responsible for their code in production, and PagerDuty's new functionalities are aimed at reducing the complexity of incident resolution through features like alert filtering, customizable workflows via new APIs, and bespoke incident actions. For ITOps, the focus is on automating issue detection, improving agile collaboration, and integrating with tools like JIRA and ITSM systems to ensure efficient response to incidents. These enhancements aim to reduce chaos and burnout, drive faster resolution, and enable teams to focus more on innovation by learning from and preventing recurring issues. The webinar promises live demos, use cases, and best practices, encouraging participants to register and engage with the new tools and techniques.
Jun 13, 2017 591 words in the original blog post.
Organizations can significantly enhance their incident response capabilities by training a diverse range of team members, beyond just senior technical leads, to serve as Incident Commanders, emphasizing the importance of soft skills over technical expertise. At PagerDuty, hands-on training and a well-defined process have enabled even junior team members to effectively lead incident responses, with key skills including directive communication, time management, and active listening. The training program involves office hours, shadowing, and role-playing exercises that gradually build confidence and familiarity with the process, ultimately reducing on-call load and burnout. By fostering a supportive community and offering resources like open-source documentation and webinars, organizations can ensure a robust and responsive incident command structure that benefits both customers and team morale.
Jun 08, 2017 1,029 words in the original blog post.
Avoiding the complexities of distributed systems is often the best approach when building software, as they inherently add complexity and can hinder productivity. By keeping processes on a single node and in memory, like the LMAX disruptor used in high-performance trading platforms, developers can achieve significant speed and efficiency. However, if high availability is necessary, it's crucial to challenge assumptions and requirements, considering whether such features are essential from the outset. When scaling becomes inevitable, minimal coordination and configuration-based approaches are recommended over transaction-level coordination to reduce brittleness. Leveraging existing tools and systems can help address common distributed system challenges, allowing developers to focus on unique business problems rather than reinventing infrastructure. By adopting strategies like Command/Query separation and Event Sourcing, developers can achieve a balance between local and distributed solutions, minimizing coordination and enhancing maintainability while using redundant storage effectively.
Jun 07, 2017 1,751 words in the original blog post.
A year ago, a technical glitch at Citi disrupted Costco Anywhere cards and ATMs, leading to widespread complaints and highlighting the need for improved incident management. Such large-scale incidents, termed "tire fires," typically demand a coordinated response involving leadership, technical teams, and external communications. Instead of focusing solely on blame through root cause analyses, organizations are encouraged to adopt a portfolio approach that assesses current investments in DevOps and support tools, allowing for strategic reallocations to enhance incident resolution. Tools like ServiceNow, PagerDuty, and Slack are crucial for rapid response and communication but require proper integration and process definitions for effective use. Metrics for incident handling, such as priority assignment, communication effectiveness, and customer satisfaction, are essential for ongoing improvement and should be communicated in clear language accessible to both technical and business stakeholders. Emphasizing defined processes and avoiding catch-all categories in evaluations can lead to more effective incident resolution and prevention strategies, ultimately aligning technical efforts with business outcomes.
Jun 06, 2017 1,100 words in the original blog post.
Business teams, like those in Business Operations and Business Intelligence, often face high volumes of urgent requests and critical projects, making traditional project management methods cumbersome. However, adopting a lightweight agile process can help these teams manage chaos effectively. PagerDuty’s business teams have implemented such a process, focusing on four main principles: transparency, prioritization, breaking down work, and measuring progress. By visualizing all current and planned work in a task management tool, teams achieve transparency, which helps in prioritizing tasks by their impact. Breaking down projects into smaller tasks allows for concurrent work and better progress measurement, facilitating reliable predictions of project completion times. Weekly meetings further enhance communication, alignment on priorities, and help in tracking progress, enabling teams to adjust plans as necessary. This agile approach allows business teams to focus on valuable work without requiring professional project managers, ultimately maximizing efficiency, transparency, and predictability.
Jun 01, 2017 664 words in the original blog post.