GitLab.com outage on 2015-09-01
Blog post from GitLab
GitLab.com experienced an hour-long outage due to what initially appeared to be a filesystem corruption issue on their NFS server, which slowed down the entire platform. Upon investigation, error messages from the server (specifically, ext4 errors) had been occurring regularly since the filesystem was first mounted, and were not a new or escalating issue as initially feared. The team opted to abort the filesystem check and restored GitLab.com to normal operation, while acknowledging the need to move data off the ext4 filesystem due to size limitations. Despite doubling the number of servers to handle traffic after a similar issue the previous week, the root cause of the NFS slowdowns remains unclear, with no direct correlation found between ext4 errors and NFS troubles. During this incident, communication with users was lacking, highlighting a need for better crisis communication practices from the operations team.
No tracked trend matches for this post yet.
Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.