August 2017 Summaries
17 posts from PagerDuty
Filter
Month:
Year:
Post Summaries
Back to Blog
PagerDuty's Summit will feature a chaos engineering-inspired Breakathon at their San Francisco office, where participants will form teams of 3-5 people to tackle infrastructure failures within a four-hour timeframe, reflecting the company's commitment to humane on-call schedules. Influenced by initiatives like Stripe's CTF and Gremlin's conversations, the event encourages creative solutions to various challenges, allowing anyone with a device capable of loading web pages and using SSH to participate. The Breakathon, limited to 50 participants, offers prizes including iPads for the winning team and Amazon gift certificates for runners-up, alongside door prizes and a raffle for Web Summit passes. Participants will enjoy a happy hour with PagerDuty engineers following the event, providing an opportunity for networking and relaxation.
Aug 31, 2017
498 words in the original blog post.
PagerDuty is launching a new bi-weekly training program called PagerDuty 101, designed to help new users understand the platform's best practices for configuration and incident response. These 60-minute sessions, conducted twice a month, are aimed at incident responders, administrators, managers, and account owners, providing them with the opportunity to learn how to set up their accounts effectively, manage user profiles, and respond to incidents using the system's features. The training covers essential topics such as inviting users, creating schedules, setting up escalation policies, configuring services, and utilizing API access keys, with a live Q&A session to address any questions. This initiative seeks to ensure that all users, including those who have recently purchased, been added to, or are trialing PagerDuty, can maximize the platform's capabilities for incident management.
Aug 30, 2017
247 words in the original blog post.
In the digital era, effective communication during crises has become increasingly vital for businesses, impacting revenue and brand equity. This pressure extends beyond technical teams to those responsible for brand reputation and customer experience, necessitating new workflows and best practices, as demonstrated by PagerDuty. At PagerDuty, the communications team is integrated into the incident response process, working closely with support, sales, legal, and leadership teams to ensure proactive and accurate communication with customers during incidents. This involves being on-call, coordinating with support to understand incidents, and maintaining real-time updates to manage customer perceptions effectively. After resolving incidents, a post-mortem analysis is conducted to improve future responses and rebuild customer trust. The approach emphasizes a collaborative, cross-functional effort to maintain digital operations excellence and deliver exceptional customer experiences, a topic further explored at the PagerDuty Summit.
Aug 29, 2017
903 words in the original blog post.
The Community team at PagerDuty, despite being newly formed as of June, is preparing for an engaging presence at the PagerDuty Summit in San Francisco. On September 6th, they will host a Breakathon at their headquarters, followed by a day at the Community Lounge on the expo floor on September 7th, where they will be available to meet attendees and answer questions. The Lounge will feature demos, hands-on activities, and live testing sessions with the User Experience team, showcasing upcoming features such as schedule management, response automation, event rules, and custom actions for their web and mobile apps. Additionally, the App Lab will offer hands-on sessions from the Product team to demonstrate how to build on the PagerDuty platform, with more surprises promised for attendees.
Aug 28, 2017
352 words in the original blog post.
The PagerDuty Summit, set for September 7th, promises an engaging experience for attendees with a lineup that includes PagerDuty University, a Breakathon, over 15 speakers, two tracks, 14 sessions, and two happy hours. The event aims to provide valuable insights from thought leaders and industry experts in the DevOps space, focusing on transforming operations and managing business responses to issues of any scale. In collaboration with Hyatt, PagerDuty is also offering attendees a chance to win a two-night stay at the Hyatt Regency Resort, Spa, and Casino in Incline Village, worth $1,000, by simply checking in at the summit registration desk before 5pm. The winner will be announced at the after-party at Pier 27, adding an extra layer of excitement to the event.
Aug 24, 2017
269 words in the original blog post.
Achieving operational maturity in IT systems is essential for organizations, encompassing consistency, reliability, resilience, and sophisticated management, design, and operation. Models, such as those from Gartner and Microsoft, define stages of operational maturity, beginning with chaotic, reactive approaches and evolving towards proactive, strategic partnerships with management that leverage IT as a competitive advantage. An application-centric approach, focusing beyond basic functionality to enhance user experience through comprehensive digital services, is pivotal in this transition. Organizations can progress from survival-focused operations to mature systems by acknowledging their current state, stabilizing infrastructure, and fostering cross-functional collaboration, ultimately transforming IT from a cost center to a source of innovation and value.
Aug 23, 2017
1,002 words in the original blog post.
PagerDuty has effectively streamlined its Customer Success team's operations by integrating Zendesk for ticketing and utilizing PagerDuty for managing on-call rotations, ensuring timely responses to customer requests and maintaining service level agreements (SLAs). The implementation of on-call shifts and escalation policies has distributed workload fairly among team members and improved response times, which has been beneficial for both customers and the Sales team. Additionally, PagerDuty employs the Net Promoter Survey (NPS) to gather customer feedback, using the insights to enhance their services, shape the product roadmap, and maintain customer satisfaction. The company has launched the PagerDuty Community to foster knowledge sharing among users and continues to iterate on its processes to ensure it meets customer expectations. PagerDuty is expanding its reach beyond IT and DevOps to Digital Operations Management, with plans to showcase its success stories at the upcoming PagerDuty Summit.
Aug 22, 2017
1,111 words in the original blog post.
Customers remain loyal to companies that share their values, making it crucial for organizations to maintain a strong brand image, especially during unexpected events. In today's digital landscape, where brands are constantly under public scrutiny, effectively managing crises involves not only addressing technical and security aspects but also integrating corporate communications early in the incident response. By doing so, companies can better control the narrative, mitigate potential damage, and maintain customer trust. The role of corporate communications is pivotal, as they prepare and execute plans to manage public perception, helping to reshape the brand once a crisis has passed. As digital channels proliferate, the expectations for seamless communication and swift, transparent incident management are rising, underscoring the importance of involving communications teams at the onset of a crisis to ensure the brand's resilience and positive public reception.
Aug 21, 2017
537 words in the original blog post.
Attending a San Francisco Giants baseball game, the narrator experienced the unexpected stress of receiving a PagerDuty alert, prompting them to leave the excitement of the game to join a conference call in a stadium bathroom. Despite feeling unprepared and flustered, as they were only shadowing and lacked necessary equipment, the incident was resolved quickly by the on-call engineer, highlighting the effectiveness of PagerDuty's platform in managing technical issues as a collaborative effort rather than an individual burden. This experience led to a realization that being on-call is akin to a team sport, where success depends on coordinated efforts and communication, reducing the pressure on any single individual. The game’s dramatic turnaround, with the Giants winning in the bottom of the ninth inning, served as an apt metaphor for the collaborative nature of managing on-call responsibilities, emphasizing how PagerDuty transforms the process into a collective endeavor, ensuring smoother customer experiences and less stressful on-call duties.
Aug 18, 2017
1,113 words in the original blog post.
PagerDuty's journey to developing an effective incident response process highlights the importance of structured improvement and collaboration over time. Initially, the company faced chaos with its rudimentary approach of alerting all personnel simultaneously, leading to uncoordinated efforts and confusion. By refining communication through a shared vocabulary and adopting Incident Command System-styled roles, PagerDuty significantly enhanced the efficiency of their response, reducing both the time taken and customer impact. The company also devised strategies to address common pitfalls, such as removing disruptive participants from calls, ensuring only essential personnel were involved. This evolution in incident management underscores the necessity for companies to prioritize and systematically refine their processes instead of relying on informal knowledge transmission, aiming for a well-prepared, comprehensive, and humane approach.
Aug 17, 2017
406 words in the original blog post.
Financial institutions face significant challenges in managing the consequences of security breaches, as they are prime targets for cybercriminals due to the valuable data and assets they hold. The Federal Deposit Insurance Corp. (FDIC) has established minimum requirements for incident response, emphasizing the importance of having a well-prepared and effective plan. Despite this, many financial organizations still react haphazardly, which can lead to wasted time and a perception of inadequate security measures. Regulators now evaluate IT security within the broader risk management standards, holding institutions accountable for both prevention and response effectiveness. A robust incident response plan is crucial, encompassing IT resolution, legal and regulatory engagement, and customer communication. Such a framework ensures that all departments have shared visibility and can mitigate breach impacts efficiently. Best practices require leadership, regular training, and clear processes to build trust and prevent loss of customer confidence. Financial institutions are encouraged to use open-sourced incident response documentation and solution briefs to establish effective workflows and processes.
Aug 16, 2017
681 words in the original blog post.
Incident management in regulated industries is particularly challenging, as it involves ensuring compliance with strict regulations that go beyond typical concerns like downtime and security breaches to include any event that could lead to non-compliance. This can encompass a wide range of incidents, such as contamination in a water supply or failure of critical systems in a hospital, where regulatory compliance is crucial. The consequences of non-compliance can be severe, including fines, legal action, loss of licenses, reputational damage, and even criminal charges. Effective incident management in these industries involves preventive measures, adherence to industry-specific guidelines and regulations, and prioritization of compliance-related incidents. Organizations need to identify sensitive systems, prevent potential failures, and have a well-coordinated incident response team ready to act swiftly to mitigate risks. The emphasis is on prevention and preparedness, as the cost of a major incident can far exceed the expenses incurred in maintaining robust incident management practices.
Aug 15, 2017
1,111 words in the original blog post.
PagerDuty has launched Dynamic Notifications, allowing users to adjust how they are notified based on the content of incoming alerts, thus eliminating the need for separate services for different urgencies. This functionality uses rules-based automation to ensure alerts are delivered with the appropriate urgency, enhancing accuracy and efficiency. For instance, info and warning alerts can trigger low-urgency notifications, while error and critical alerts prompt high-urgency responses, with escalation as needed. The system is supported by an Event Rules Engine, offering further customization for incident notification urgency, even if upstream tools cannot send specific severities. Importantly, Dynamic Notifications can be tailored to specific time windows through the Support Hours feature, enhancing flexibility. Available at no extra cost to customers on Standard and above plans, this feature aims to streamline alert management and improve service-based reporting and analytics.
Aug 15, 2017
350 words in the original blog post.
Ellen, a Computer Science and Psychology student at Swarthmore College, reflects on her enriching summer internship at PagerDuty as a DataDuty intern, where she contributed to data infrastructure projects and learned about data engineering beyond typical software roles. Her experience featured engaging in initiatives like Hack Day, where she developed a Slack bot that won a company award, and attending the Girls in Tech Catalyst Conference, which highlighted PagerDuty's commitment to diversity and inclusivity. Ellen appreciated the supportive environment that encouraged open communication and learning, allowing her to shadow various departments and gain insights into the business side of the company. This internship not only allowed her to undertake significant projects but also fostered relationships with mentors and peers, leaving her with valuable career and life lessons as she returns to college, armed with new goals and a broadened perspective on her future career path.
Aug 10, 2017
1,377 words in the original blog post.
PagerDuty is launching its first PagerDuty University at the PagerDuty Summit to address customer demands for more effective management of digital operations and incident response. The training workshop aims to familiarize users with newer capabilities to enhance efficiency in managing digital operations and handling major incidents. It is divided into two main tracks: "Owning Incident Response" and "Improving the On-Call Life." The first track emphasizes a calm, systematic approach to incident management, drawing on best practices and simulation exercises, while the second track focuses on maximizing the use of PagerDuty’s platform, exploring its integrations, and managing alerts effectively. The initiative seeks to bring together both new and experienced PagerDuty users to share best practices and address common challenges, with the ultimate goal of improving the on-call experience and leveraging PagerDuty’s tools more effectively.
Aug 09, 2017
792 words in the original blog post.
Major brands often face significant infrastructure failures during high-demand periods like Black Friday, raising concerns for smaller companies about maintaining uptime during normal operations. By adopting robust incident management procedures, smaller retail teams can effectively mitigate disruptions. Retailers must focus on maintaining customer-facing websites, backend systems, point-of-sale systems, and IoT assets to ensure business continuity. Effective strategies include maximizing visibility across infrastructure, deploying flexible monitoring solutions, responding in real time, and facilitating seamless communication through collaborative tools like ChatOps. While the complete elimination of downtime threats is unlikely, modern monitoring and incident management solutions are crucial in preventing service failures.
Aug 08, 2017
849 words in the original blog post.
Vinh Tran's internship at PagerDuty as a Software Engineer Intern on the Platform team has been a period of rapid learning and professional growth, marked by challenging coding tasks, active participation in team meetings, and opportunities to contribute to real projects. The onboarding process was well-structured with a detailed checklist that allowed Tran to quickly become productive and start contributing code to the production environment. The supportive company culture, characterized by trust and encouragement from colleagues and managers, has bolstered his confidence and inspired a shift in mindset from doubt to determination. Tran has come to appreciate the importance of code reviews, which have taught him to scrutinize and refine his code while learning from peers' submissions. The agile workflow at PagerDuty, with its emphasis on continuous improvement through sprints and retrospectives, has provided a framework for Tran to enhance his skills. Furthermore, the company's fast-growth environment and events like HackDay have fostered innovation and creativity, resulting in Tran's team winning an award for Best Product Enhancement. This experience has not only honed his technical abilities but also encouraged growth in communication and professional development, preparing him for future roles in the tech industry.
Aug 03, 2017
1,021 words in the original blog post.