May 2015 Summaries
5 posts from PagerDuty
Filter
Month:
Year:
Post Summaries
Back to Blog
Five engineers and product managers from PagerDuty are set to speak at the Velocity Santa Clara conference, covering a range of topics pertinent to both engineering and business. Evan Gilman will discuss integrating engineering and business teams in a collaborative manner, while Doug Barth will address securing network traffic using an IPSec mesh network. Amanda Folson is participating in a panel on managing burnout in the tech industry, focusing on the challenges of workload and maintaining mental health. Additionally, Dave Cliffe and Arup Chakrabarti will present on managing Black Swan incidents, offering insights into incident impact and response through engaging metaphors. The talks are scheduled for late May 2015, promising valuable insights into modern engineering and business practices.
May 22, 2015
311 words in the original blog post.
CloudMonix, a successor to AzureWatch, is a comprehensive cloud services monitoring application designed to provide deep insights, instant notifications, and automatic problem resolution for complex cloud environments, primarily focusing on Microsoft Azure. It offers cloud administrators a range of tools such as live dashboards, key performance metrics, instant alerts, self-healing actions, auto-scaling, and API integration to monitor and manage Azure services like virtual machines, SQL databases, and storage services. The platform's integration with PagerDuty enhances incident management by allowing users to receive and manage CloudMonix alerts through PagerDuty's system, facilitating quick escalation and resolution of system issues. This collaboration ensures that cloud administrators maintain better control over their IT infrastructure, allowing for prompt responses to system-wide issues. CloudMonix plans to expand its support to other cloud platforms, such as Amazon Web Services and OpenStack, further broadening its utility for cloud system administrators.
May 19, 2015
630 words in the original blog post.
PagerDuty's new feature, Rich Incidents, is designed to enhance the incident response process by providing real-time data and seamless communication tools directly from an alert. This feature allows incident responders to quickly access conference bridges, chat rooms, or runbooks, enabling them to collaborate efficiently and leverage institutional knowledge. It includes the integration of embedded graphs and images from partners like Datadog and Ghost Inspector, offering immediate visual context and historical data to better understand the severity and scope of incidents. Datadog's integration focuses on sharing real-time metrics, while Ghost Inspector provides automated browser testing alerts complete with images and video run-throughs. Users can also create custom graphs using the API for tailored visualizations in alerts, enhancing the ability to respond swiftly and effectively to critical incidents.
May 13, 2015
563 words in the original blog post.
ZooKeeper, a prominent open-source project known for enabling distributed coordination, encountered significant reliability issues at PagerDuty due to a confluence of bugs in both ZooKeeper and the Linux kernel. These issues resulted in random cluster-wide lockups, largely stemming from two ZooKeeper bugs related to client session overloads and unhandled exceptions in critical threads, and two kernel-related bugs involving TCP payload corruption and checksum validation failures under specific conditions. The investigation revealed that TCP payload corruption was linked to the use of IPSec in Transport Mode combined with certain versions of the Linux kernel and Xen virtualization, which allowed corrupted packets to bypass validation. Further complicating matters, the aesni-intel kernel module was implicated in the corruption during AES encryption. Despite arduous troubleshooting efforts, including downgrading affected systems and blacklisting problematic modules, a definitive fix remains elusive, although workarounds have been implemented to mitigate the issues. The investigation underscores the complex interplay between software components and the challenges in maintaining high reliability in distributed systems.
May 07, 2015
3,009 words in the original blog post.
Boundary has integrated its real-time, second-by-second IT system monitoring solution with PagerDuty to enhance incident resolution and communication for IT and DevOps teams. This integration allows for seamless triggering and resolution notifications, improving the speed and efficiency of addressing infrastructure incidents. Boundary's monitoring service is designed for ease of use with a simple setup, intuitive UI, and compatibility with over 20 server operating systems, while offering customizable dashboards and alarms for various popular technologies. The integration with PagerDuty, a popular service known for incident visibility and collaboration, aligns with Boundary's commitment to open communication and streamlined operations. The updated integration simplifies the process of connecting Boundary alarms to PagerDuty, ensuring that issues are quickly identified and addressed by the appropriate on-call personnel, thereby enhancing the problem-solving capabilities for joint customers.
May 05, 2015
596 words in the original blog post.