How we improved on-call life by reducing pager noise
Blog post from GitLab
GitLab.com has implemented a system to monitor its services using Service Level Indicators (SLIs) and Service Level Objectives (SLOs) to promptly address issues before they impact users. However, the previous setup led to excessive paging of the Site Reliability Engineering (SRE) team during service-wide or site-wide outages, causing stress and distraction. To address this, GitLab introduced a more efficient alert management system using Prometheus and Alertmanager, which groups alerts by service and incorporates service dependencies. This change reduces the number of pages the SRE team receives, allowing them to focus on resolving issues more effectively. The system now ensures that if a foundational service like the database degrades, it prevents cascading alerts for dependent services, thus improving the quality of life for on-call engineers by consolidating alerts and decreasing unnecessary notifications. This approach has already resulted in a significant reduction in the number of pages received during outages.
No tracked trend matches for this post yet.
Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.