Home / Companies / GitLab / Blog / Post Details
Content Deep Dive

GitLab.com outage on 2015-05-29

Blog post from GitLab

Post Details
Company
Date Published
Author
Jacob Vosmaer
Word Count
635
Company Posts That Month
9
Language
English
Hacker News Points
-
Post removed?
No
Summary

GitLab.com experienced an outage on May 29, 2015, due to a backup script malfunction that caused the filesystem on the backend server to remain frozen, leading to a service disruption that lasted approximately 94 minutes. The outage was compounded by the recent infrastructure upgrade, which added complexity to the system, and by insufficient training and documentation for on-call engineers. The recovery process involved diagnosing issues with the NFS share and backend server, ultimately requiring a reboot to restore functionality. In response, GitLab removed the freeze/unfreeze steps from the backup script, implemented a secondary backup strategy for SQL data, and prioritized training through regular operations drills to enhance the preparedness of their engineering team.

Trends Found in this Post

No tracked trend matches for this post yet.

Use This Data

Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.