March 2017 Summaries
16 posts from PagerDuty
Filter
Month:
Year:
Post Summaries
Back to Blog
PagerDuty has introduced the "OnCallClub," a loyalty program designed for on-call developers, system administrators, and others with root access, who often face early morning wake-up calls. This program allows participants to convert their incident management efforts into rewards through a system of PagerPoints™, with loyalty tiers such as Silver, Premier, and Sapphire, along with secret levels offering perks like custom sleepwear and on-demand beverages. Despite its innovative concept, the program has been met with skepticism from some inaugural members, who humorously question its practicality and express relief that it is not a real initiative.
Mar 31, 2017
234 words in the original blog post.
Fear of failure can significantly impact the morale and productivity of development and operations teams, but an effective incident management solution can alleviate this fear and empower teams. By centralizing alerts and providing relevant data, incident management reduces stress during incidents, enabling faster issue resolution and fostering happier teams. It also brings predictability when adding new features, as teams gain a deeper understanding of system functions and recurring issues, leading to informed feedback and increased confidence. This alignment between development and operations creates a unified platform that enhances user experiences by providing a realistic view of uptime, aligning team goals with business objectives. Additionally, improved communication across teams reduces bottlenecks, allowing more time for development and quicker deployment, which in turn fosters a sense of empowerment and ownership among team members. By reducing reliance on higher-ups for decision-making and encouraging accountability, incident management enhances team confidence, innovation, and overall morale, leading to successful outcomes.
Mar 30, 2017
816 words in the original blog post.
High-growth companies, characterized by rapid sales increases and expanding employee numbers, face unique challenges and opportunities that can either propel them forward or lead to their downfall. To navigate this critical phase, companies should set clear goals and foster cross-departmental collaboration to ensure alignment and responsible growth. Recognizing and rewarding achievements in real-time strengthens team morale and commitment. Developing a scalable on-boarding process is crucial to integrate new employees efficiently, promoting a culture of learning and communication that accelerates productivity and reduces attrition. Maintaining strict hiring standards is vital to safeguard company culture and ensure that only top talent is recruited, as hiring mistakes can significantly impact high-growth startups. Ultimately, while high-growth startups present exciting opportunities for innovation and market disruption, it is essential to prioritize intelligent, sustainable growth over rapid expansion.
Mar 29, 2017
1,025 words in the original blog post.
Incident management is crucial for modern IT operations teams, yet scaling it can introduce challenges due to the increasing complexity of monitoring a growing landscape of devices, applications, and systems. As teams expand and adopt hybrid IT models, onboarding new engineers and implementing effective notification policies become more complex. A common scenario involves integrating new IT environments after a business acquisition, which often entails dealing with different tech stacks and incident management tools. Key strategies for effective scaling include identifying areas of growth, ensuring comprehensive monitoring tool coverage across the stack, and implementing systems to centralize, normalize, and deduplicate monitoring data for actionable insights. Reducing noise through effective data routing and thresholding is essential to prevent alert fatigue, while a robust incident management platform helps unify alerts, supports team growth, and fosters accountability and collaboration. As IT operations evolve towards hybrid and agile frameworks, scaling incident management is essential to meet user demands for reliable data access and to mitigate the increasing stakes of downtime.
Mar 28, 2017
738 words in the original blog post.
Incident response bottlenecks are critical challenges that can hinder the effectiveness of on-call teams and negatively impact customers, necessitating strategies to minimize them. Key goals of incident response include preventing incidents, confining damage, and resolving issues swiftly. Bottlenecks often arise from inadequate prioritization, alert fatigue, insufficient training, and lack of preparation for new rollouts. Prioritization is essential for focusing on high-impact incidents, while automated systems can help filter alert noise and direct alerts to the right teams, reducing alert fatigue. Effective training and documentation, such as runbooks, can mitigate the impact of inexperienced team members. Additionally, preparedness for major rollouts through limited deployments can prevent resource depletion during high-priority alert storms. While other bottlenecks may exist, addressing these core issues can significantly enhance incident response efficiency.
Mar 23, 2017
1,067 words in the original blog post.
Joe Sexton, a new member of PagerDuty’s Executive Advisory Board, shares insights on career advancement, emphasizing that it doesn't necessarily require decades of experience to fast-track one's career. He highlights three key areas: negotiation, under-promising and over-delivering, and networking. In negotiation, preparation and understanding customer needs are crucial, while maintaining long-term relationships and being creative with solutions can lead to better outcomes. Under-promising and over-delivering can lead to increased responsibility and career growth, as it builds trust and confidence among colleagues. Networking is essential, regardless of personality type, and involves maintaining connections, asking for help, and offering assistance to others. Sexton concludes that these strategies, along with leveraging individual skills and choosing supportive companies, can significantly accelerate career progression, encouraging readers to explore opportunities at PagerDuty.
Mar 22, 2017
1,150 words in the original blog post.
Having the right tools and procedures in place before a major outage is crucial for effectively managing such incidents. Proper internal communication tools like Slack or HipChat can facilitate real-time updates and disseminate information efficiently, unlike traditional email. Application performance and infrastructure monitoring technologies, such as New Relic and AWS CloudWatch, are essential for identifying issues before customers do and providing valuable insights during an outage, while solutions like PagerDuty integrate these insights for a comprehensive view. Status updates via platforms like Twitter or statuspage.io help maintain customer trust and transparency during disruptions, and tools like ZenDesk are vital for managing support requests and documenting issues. Procedure tracking, including pre-documented scenarios and conducting mock outages, is key to anticipating potential problems and improving response strategies. The effectiveness of these tools and strategies depends on their proper configuration and understanding before an incident occurs, highlighting the importance of communication with stakeholders during such events.
Mar 21, 2017
888 words in the original blog post.
On-call engineers play a pivotal role in incident management, particularly in determining whether an incident escalates or is efficiently resolved. As organizations grow, establishing a structured process for these engineers becomes crucial, regardless of company size. Key aspects include a rapid first response, understanding system functionality, automatic scheduling for fair rotation, and having backup engineers to ensure no incidents are overlooked. Proper training, as well as tools like checklists and flowcharts, aid engineers in swiftly managing incidents, which involve identifying, logging, categorizing, and prioritizing issues. Effective communication, often facilitated by platforms such as PagerDuty, is vital for mobilizing the right personnel quickly. Troubleshooting should commence immediately, even before the entire team is assembled, to optimize response time and minimize business impact. Robust planning and resource management in these processes allow teams to focus more on innovation rather than problem-solving.
Mar 16, 2017
857 words in the original blog post.
In incident management, suppression is a crucial technique used to manage the overwhelming volume of alerts generated by modern infrastructure, ensuring that high-priority alerts receive the necessary attention while preventing alert fatigue among admins. Rather than permanently deleting data, suppression temporarily withholds certain alerts from appearing on priority dashboards, allowing admins to focus on actionable incidents. This approach can be nuanced, with configurations allowing for alerts to be suppressed or reported based on criteria like frequency, time of day, or device type. Importantly, suppression does not entail data loss; suppressed alerts are still recorded and can be reviewed as needed, contributing to historical data analysis and allowing for better tuning of alerting thresholds. This method enhances the efficiency of incident management by reducing noise without sacrificing the visibility of critical infrastructure events.
Mar 15, 2017
781 words in the original blog post.
Technical debt, which refers to inefficiencies or imperfections in software code or architecture that accumulate over time, can be effectively monitored and managed using incident management tools without requiring additional resources. These tools, such as PagerDuty, allow for ongoing monitoring of IT infrastructure, automatically providing insights into technical debt by tracking metrics like the number of alerts, mean time to resolution (MTTR), and escalation rates. An increase in alerts or escalations, or a high MTTR, can indicate inefficiencies or complex code, helping identify areas where technical debt is accumulating. By utilizing existing monitoring systems, organizations can proactively address technical debt, improving code efficiency and operational performance without additional investment.
Mar 14, 2017
1,042 words in the original blog post.
Downtime can be costly for businesses, with expenses arising from lost revenue, wasted employee productivity, and unused resources, potentially leading to a loss of customer trust. The primary causes of outages include human error, third-party service failures, and unpredictable events, with solutions focusing on checks and balances such as code reviews, unit tests, and quality assurance. Additionally, tools like Netflix's Chaos Monkey and PagerDuty's incident management system help organizations prepare for and manage service disruptions. Effective communication with customers during outages is crucial to maintaining trust, and tools such as StatusPage can provide transparency. Establishing on-call rotations ensures that there are always personnel available to address issues promptly while minimizing disruption to employees' personal lives. Investing in these resources and processes can significantly reduce downtime impacts, underscoring the importance of preparedness in maintaining business continuity.
Mar 09, 2017
954 words in the original blog post.
International Women’s Day serves as a global celebration of women's achievements across various spheres, emphasizing unity, reflection, and advocacy. The author reflects on their career experiences with inspirational female leaders, particularly one manager who, despite her reserved demeanor, demonstrated exceptional leadership qualities such as preparedness, humility, and an ability to balance practical and inspirational approaches. Her leadership style, characterized by respect, courtesy, and authenticity, provided a model that transcended traditional gender norms in leadership. At PagerDuty, where the author currently works, diversity and inclusivity are core values, supporting professional growth for all employees. The narrative underscores the importance of strong female role models and inclusive work environments, expressing gratitude that these will be available to inspire future generations, including the author's daughter, as she embarks on her career in engineering.
Mar 08, 2017
492 words in the original blog post.
Jennifer Tejada, CEO of PagerDuty, reflects on the significance of International Women's Day, emphasizing its importance beyond just a holiday. She acknowledges the "A Day Without Women" movement, which highlights women's rights and equal social justice, expressing pride in working for an inclusive and diverse company. Tejada shares her personal plans to celebrate the day with her daughter and highlights the importance of discussing women's voices and opportunities for the future. She encourages employees to support women in their own ways while maintaining consideration for colleagues and customers. Tejada reaffirms her commitment to fostering an inclusive environment at PagerDuty, advocating for diversity as a key factor for success and encouraging personal activism that supports women.
Mar 08, 2017
551 words in the original blog post.
Mobile incident management is essential in today's fast-paced environment, providing real-time alerts and resolution capabilities from any device, ensuring incidents can be managed effectively regardless of location. This approach leverages smart devices to fill critical gaps in incident management, allowing teams to resolve issues on-the-go with features like alerts, timelines, and easy integration with other apps. Particularly beneficial for DevOps teams, mobile incident management enhances collaboration across development, QA, and operations, improving response times and encouraging better code production. On-call engineers benefit significantly from mobile apps, which streamline incident triage, prioritization, and resolution, minimizing unnecessary escalations and allowing for efficient teamwork. Tools like PagerDuty's mobile app offer centralized dashboards and customizable access controls, making mobile incident management an indispensable part of the workflow, ensuring issues are addressed promptly and efficiently.
Mar 07, 2017
855 words in the original blog post.
PagerDuty hosted a webinar with industry leaders to discuss the evolving practices and future of DevOps, examining whether it has become mainstream and how enterprises can adopt these practices. Panelists included leaders from Datadog, Threat Stack, AppDynamics, and PagerDuty, who shared insights on how DevOps is perceived and implemented in various organizational contexts. The consensus was that while DevOps is widely recognized, its true adoption varies, particularly between startups and larger enterprises. Key challenges include cultural alignment, sharing, and breaking down silos to enable effective implementation. The panelists emphasized that while tools are important, success in DevOps relies heavily on cultural change and organizational alignment, which are essential for driving business value. The discussion also highlighted that DevOps is becoming increasingly necessary for enterprises, with a focus on rapid software releases and aligning IT with business objectives.
Mar 02, 2017
1,168 words in the original blog post.
Embarking on a journey with PagerDuty a year ago, the author reflects on the company's transformation from a small startup into a leading Digital Operations Management Platform under the leadership of CEO Jennifer Tejada. Initially joining as an Enterprise Business Representative, the author has witnessed the company's culture of inclusivity and innovation, which emphasizes maintaining company values, employee happiness, and fostering a DevOps community. Attending events like DevOps Days Minneapolis has highlighted the impact PagerDuty has on its customers, particularly in helping engineers achieve a better work-life balance. The past year has been marked by significant growth, including new leadership and processes, and the author is now transitioning into the role of Business Development Manager for Executive Briefings, expressing excitement about the company's future and inviting others to join in making an impact.
Mar 01, 2017
536 words in the original blog post.