How we share SLIs across engineering departments
Blog post from GitLab
GitLab employs a collaborative approach to maintain the availability of GitLab.com, using Service Level Indicators (SLIs) to monitor performance across different development groups, each with its own dashboard. The target Service Level Objective (SLO) for GitLab.com's availability is set at 99.95%, and if a group's performance falls below this threshold, efforts are made to identify and resolve the issues affecting feature reliability. The GitLab infrastructure is divided into multiple services that run the same Rails application but handle different types of traffic, requiring isolated and aggregate monitoring to accurately assess performance. This monitoring is facilitated by Grafana dashboards, which display error rates and generate multi-window, multi-burn-rate alerts based on Google's SRE practices. These alerts help bring attention to issues that may not be apparent in larger service aggregations, particularly for features with lower traffic, allowing teams to prioritize necessary improvements. Future discussions will focus on the development and integration of these monitoring tools into GitLab's product prioritization process.
No tracked trend matches for this post yet.
Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.