Home / Companies / PagerDuty / Blog / April 2011

April 2011 Summaries

4 posts from PagerDuty

Filter
Month: Year:
Post Summaries Back to Blog
PagerDuty recently showcased its offerings at the Under The Radar event, where it was honored with both the Best In Show’s Audience Choice Award and the Developer Tools’ Audience Choice Award, thanks to the support of the audience. The presentation that earned these accolades is available for viewing through a video, with accompanying slides provided below it, offering valuable insights into PagerDuty's capabilities.
Apr 29, 2011 50 words in the original blog post.
Amazon experienced significant issues with its cloud infrastructure services, EC2, EBS, and RDS, which began around 1 am Pacific Time, affecting many internet sites and services. This outage has been a critical moment for PagerDuty, a platform used to alert operations staff about system problems, as it saw a substantial increase in alerts and events being processed. During the outage, PagerDuty routed notifications to 36% of its customer base, indicating that a significant portion of its clients experienced issues severe enough to require waking up sysadmins or engineers. Additionally, more than 10% of all operations staff across PagerDuty's customer base were alerted, suggesting extensive disruptions. The company noted that while they typically handle initial alerts, the AWS problems likely triggered comprehensive emergency responses within affected organizations. The situation led to a spike in both outgoing alerts and incoming events to PagerDuty, reflecting the broader impact of the AWS service disruptions on internet functionality.
Apr 22, 2011 496 words in the original blog post.
The introductory post in a series on enhancing service or system availability explains key concepts such as availability, service-level agreements (SLAs), mean time between failure (MTBF), and mean time to recovery (MTTR). Availability refers to the percentage of time a system is operational, with high availability often expressed in "nines," such as 99.99%, equating to minimal annual downtime. The post criticizes the exclusion of scheduled downtime from availability calculations, arguing that downtime affects availability regardless of scheduling. SLAs define the minimum availability levels required before financial compensation is due to customers, although such refunds are typically minor compared to the broader financial impact of outages, which can harm customer trust and business reliability. MTBF and MTTR are practical measures of system reliability, with a focus on increasing MTBF by building robust systems and reducing MTTR through preparedness and efficient recovery strategies. The following posts in the series will delve into methods for decreasing MTTR.
Apr 19, 2011 637 words in the original blog post.
PagerDuty humorously announced a fictitious pivot from their successful server alert service to an innovative, albeit satirical, transportation solution called the Curated Arial Non-Orbital Navigation System (CANON) as part of an April Fools joke. The post critiques the inefficiencies of the US public transportation system, attributing the reliance on personal vehicles to the lack of viable alternatives and the over-regulation in the taxi industry. CANON is presented as an imaginative, air-based transportation network capable of efficiently serving even the least densely populated areas with quick setup and minimal physical footprint. Despite the fictional nature of the project, it is humorously claimed to have completed a successful test phase and secured government funding, with plans for a pilot in the valley and a future rollout to major US cities. While acknowledging challenges like rapid acceleration and deceleration, the announcement cleverly maintains a playful tone, reassuring that PagerDuty will continue its original mission of alerting users to server issues.
Apr 02, 2011 762 words in the original blog post.