June 2012 Summaries
4 posts from PagerDuty
Filter
Month:
Year:
Post Summaries
Back to Blog
PagerDuty has successfully migrated its primary infrastructure from the US-East region to US-West (Oregon) on AWS, completing the process without any downtime on June 19. This strategic move addresses the issue of correlated failures experienced by over 20% of its customers who operate on AWS in the US-East region, where outages previously resulted in increased load and capacity loss for PagerDuty. By relocating, PagerDuty has alleviated these risks, while still maintaining a fully redundant hot-backup on a separate provider. Looking forward, PagerDuty plans to implement a multi-data center setup and adopt a multi-node data store using Cassandra, with further details on the new system design to be shared soon.
Jun 22, 2012
228 words in the original blog post.
A new integration between Scout, a hosted server monitoring system, and PagerDuty, an incident response platform, has been announced, enabling users to receive notifications for server, database, and application issues through phone, SMS, and email. Scout, which monitors server metrics and various services such as MySQL and MongoDB, allows for easy visualization of metrics across servers. The integration process is straightforward, requiring users to connect their Scout and PagerDuty accounts via an API-based integration, after which Scout can trigger incidents in PagerDuty. Early customer feedback has been positive, highlighting the ease of setup and the convenience of auto-resolving alerts.
Jun 19, 2012
253 words in the original blog post.
On June 14, PagerDuty experienced a significant outage beginning at 8:44 pm Pacific time, resulting in 30 minutes of downtime followed by a period of high load. The application, primarily hosted on AWS in the US-East region, suffered from an AWS console failure, prompting an emergency switch to a backup provider. This emergency flip was completed by 9:14 pm, restoring operations but under high load, which was resolved by 10:03 pm. Despite the successful flip, the team identified several areas for improvement, including faster monitoring notification, better group call organization, and a more efficient flip process. To prevent future occurrences, PagerDuty plans to migrate its data center to AWS US-West, implement a three-provider setup for greater fault tolerance, and enhance its internal monitoring and communication tools. Additionally, they aim to streamline and automate the emergency flip process and conduct load testing to improve system performance during high-load scenarios.
Jun 19, 2012
1,196 words in the original blog post.
PagerDuty customers can benefit from a new partnership with Server Density, which offers a year of free server monitoring for one server to new accounts as part of a DevOps deal. Server Density is a SaaS-based tool that provides monitoring, analytics, and alerts for servers and websites, integrating seamlessly with PagerDuty. This collaboration highlights a monthly feature by Server Density, spotlighting dev-ops tools they use internally, and this month they chose PagerDuty. The service allows users to monitor various server metrics, including CPU, memory, and disk space, as well as services like Apache, Nginx, MySQL, and MongoDB. To take advantage of this offer, interested parties should email Server Density with their account URL and reference the PagerDuty DevOps deal.
Jun 12, 2012
191 words in the original blog post.