Home / Companies / PagerDuty / Blog / Post Details
Content Deep Dive

The ups and downs of Availability

Blog post from PagerDuty

Post Details
Company
Date Published
Author
John Laban
Word Count
637
Company Posts That Month
4
Language
English
Hacker News Points
-
Post removed?
No
Summary

The introductory post in a series on enhancing service or system availability explains key concepts such as availability, service-level agreements (SLAs), mean time between failure (MTBF), and mean time to recovery (MTTR). Availability refers to the percentage of time a system is operational, with high availability often expressed in "nines," such as 99.99%, equating to minimal annual downtime. The post criticizes the exclusion of scheduled downtime from availability calculations, arguing that downtime affects availability regardless of scheduling. SLAs define the minimum availability levels required before financial compensation is due to customers, although such refunds are typically minor compared to the broader financial impact of outages, which can harm customer trust and business reliability. MTBF and MTTR are practical measures of system reliability, with a focus on increasing MTBF by building robust systems and reducing MTTR through preparedness and efficient recovery strategies. The following posts in the series will delve into methods for decreasing MTTR.

Trends Found in this Post

No tracked trend matches for this post yet.

Use This Data

Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.