June 2015 Summaries
5 posts from PagerDuty
Filter
Month:
Year:
Post Summaries
Back to Blog
PagerDuty hosted an event where representatives from Dropbox, Flipboard, Splunk, and PagerDuty discussed their insights and experiences related to operational maturity. Operational maturity was defined as the ability to understand trade-offs and impacts in production environments, recognize the implications of incidents on business and employee well-being, and effectively manage crises through informed decision-making. Dropbox highlighted the importance of a structured incident response process, Flipboard emphasized the evolution of on-call and escalation policies for better employee satisfaction, Splunk focused on meaningful alerting and continuous improvement, while PagerDuty showcased their commitment to reliability through practices like "Failure Friday" and robust incident management. The event underscored the significance of operational maturity in enabling businesses to be agile, accountable, and adaptable in a rapidly changing market.
Jun 26, 2015
654 words in the original blog post.
PagerDuty decided against developing a native chat tool to maintain alignment with the DevOps philosophy of transparency and collaboration, which emphasizes having all information in one central location. A native chat client during incidents could lead to fragmented communication records and disrupt the integration of essential data, such as deployment details and scripts, crucial for effective incident management and post-mortem analysis. Additionally, the learning curve associated with a new tool during critical incidents would detract from responders' ability to efficiently address issues. Instead, PagerDuty integrates with established chat platforms like HipChat, Slack, and Flowdock, leveraging these partnerships to allow users to work with familiar tools while focusing on building a robust IT Operations Management platform. Through open-source collaboration and a well-documented REST API, PagerDuty encourages the community to create tools that enhance its functionality with existing chat clients.
Jun 18, 2015
339 words in the original blog post.
Operationally mature DevOps teams leverage data-driven metrics to enhance performance, drive cultural change, and enable quick decision-making with minimal risk. Key metrics include Time to Response, which fosters a culture of high achievement by holding team members accountable for acknowledging incidents promptly; Escalations, which help manage expectations by tracking incident responses and determining necessary adjustments; Raw Incident Count, which combats alert fatigue by ensuring that responders focus on significant alerts; and Mean Time to Resolution, which gauges operational readiness by measuring how quickly incidents are resolved. While metrics provide valuable insights into past performance, they should be used as tools to foster future improvements and not merely as a means to assign blame, emphasizing the importance of action-oriented strategies to achieve business goals.
Jun 11, 2015
788 words in the original blog post.
PagerDuty emphasizes the importance of monitoring business metrics in real-time to prevent larger, business-impacting outages, advocating for integration of these metrics into operational workflows alongside traditional system metrics like CPU usage. By doing so, companies can preemptively identify and respond to potential issues before they escalate, ensuring a more reliable and customer-focused operation. This approach is especially crucial for e-commerce and streaming services, where unexpected changes in key metrics, such as a drop in orders or stream starts, can signal significant problems. PagerDuty practices what it preaches by ensuring their alerting pipeline is robust enough to trigger immediate responses without human intervention. The emphasis is on understanding how operational activities directly contribute to business value, encouraging engineers to adopt a business-focused perspective in monitoring activities.
Jun 04, 2015
656 words in the original blog post.
Anthony Gibbons, Operations Manager at UK-based Airhead Education, shares his experience of setting up IT operations software for the startup, emphasizing the integration of cloud-based technologies despite budget constraints. Upon joining Airhead in 2014, Gibbons aimed to enhance infrastructure monitoring and notification systems using advanced tools like Microsoft SCOM, Site 24×7, and New Relic, which were surprisingly accessible to startups. Initially challenged by missed alerts and delayed updates, Gibbons discovered PagerDuty through a New Relic promotion, which streamlined alert management by integrating with existing monitoring solutions, facilitating efficient notifications and incident management. Additionally, integrating PagerDuty with StatusPage.io and HipChat improved communication with customers and incident tracking. Gibbons highlights the adaptability and continuous evolution of PagerDuty, which aligns with Airhead's philosophy of leveraging cutting-edge technologies, underscoring the advantages of modern IT operations for startups.
Jun 02, 2015
971 words in the original blog post.