Home / Companies / GitLab / Blog / Post Details
Content Deep Dive

How we improved on-call life by reducing pager noise

Blog post from GitLab

Post Details
Company
Date Published
Author
Steve Azzopardi
Word Count
940
Company Posts That Month
16
Language
English
Hacker News Points
-
Post removed?
No
Summary

GitLab.com has implemented a system to monitor its services using Service Level Indicators (SLIs) and Service Level Objectives (SLOs) to promptly address issues before they impact users. However, the previous setup led to excessive paging of the Site Reliability Engineering (SRE) team during service-wide or site-wide outages, causing stress and distraction. To address this, GitLab introduced a more efficient alert management system using Prometheus and Alertmanager, which groups alerts by service and incorporates service dependencies. This change reduces the number of pages the SRE team receives, allowing them to focus on resolving issues more effectively. The system now ensures that if a foundational service like the database degrades, it prevents cascading alerts for dependent services, thus improving the quality of life for on-call engineers by consolidating alerts and decreasing unnecessary notifications. This approach has already resulted in a significant reduction in the number of pages received during outages.

Trends Found in this Post

No tracked trend matches for this post yet.

Use This Data

Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.