DEV Community

Cover image for rm -rf on the wrong server: how GitLab lost its production database

rm -rf on the wrong server: how GitLab lost its production database

On January 31, 2017, a GitLab engineer ran rm -rf on the PostgreSQL data directory of what he thought was the replica. It was the primary. About 300 GB of GitLab.com's production database was gone within a second or two, and when the team reached for backups, none of the five they had worked. GitLab.com was down for about 18 hours and lost six hours of data for good. It is the most famous database postmortem there is, and nearly every cause in it is still sitting in someone's infrastructure today.

TL;DR

  • While re-syncing a lagging replica late at night, an engineer deleted /var/opt/gitlab/postgresql/data on db1.cluster.gitlab.com (the primary) instead of db2.cluster.gitlab.com (the replica). Of about 310 GB, 4.5 GB remained.
  • GitLab's next-day post: "out of five backup/replication techniques deployed none are working reliably or set up in the first place" (GitLab).
  • The nightly pg_dump to S3 had been failing silently because of a PostgreSQL version mismatch; the failure emails bounced.
  • The restore came from a manual snapshot taken six hours earlier for an unrelated test. Copying it back took about 18 hours at roughly 60 Mbps.
  • About 5,000 projects, 5,000 comments and 700 new user accounts were lost. Git repositories and wikis were not affected.

What does rm -rf do, and why couldn't it be undone?

rm removes files. -r (or -R) makes it recursive, so it deletes a directory and everything under it. -f forces it: no confirmation prompts, no errors for missing files. GitLab's live incident notes record the flags as rm -Rvf, where -v prints each file as it goes. The command, as documented by GitLab's incident post and the live Google Doc:

# on db1.cluster.gitlab.com, the primary
rm -Rvf /var/opt/gitlab/postgresql/data
Enter fullscreen mode Exit fullscreen mode

There is no trash can for rm. When the CEO suggested trying to undelete the files, the live doc answers: "Not possible! rm -Rvf". Another engineer asked about open file descriptors; PostgreSQL does not keep all its files open, so that did not work either. Once the data directory is unlinked, recovery means backups.

Timeline of the GitLab database incident

All times UTC, from GitLab's postmortem (with an apology from CEO Sid Sijbrandij), the Feb 1 post and the live doc:

Time What happened
Jan 31, ~17:20 An engineer takes a manual LVM snapshot of production to test pgpool-II load balancing in staging
~19:00–21:00 Database load spikes: spam snippets, a background job hard-deleting a GitLab employee a troll reported for abuse, and one user whose repository served as a CDN, with 47,000 IPs signing into one account (@gitlabstatus)
~22:00–23:00 The replica, db2, falls about 4 GB behind and stops replicating. The WAL segments it needs are already gone from the primary
~23:00 Re-sync attempts. pg_basebackup hangs with no output
~23:25–23:27 The engineer deletes the data directory on db1, notices within a second or two, and stops it. About 300 GB are gone
23:28 "We are performing emergency database maintenance, GitLab.com will be taken offline" (tweet)
Feb 1, 00:44 "We accidentally deleted production data and might have to restore from backup"
Feb 1, ~17:00 Database restored from the six-hour-old staging copy, without webhooks
Feb 1, ~18:00 Webhooks restored; "6:14pm UTC: GitLab.com is back online"
Feb 10 Postmortem published with the fix list and issue numbers

@gitlabstatus, Feb 1 2017 00:44 UTC:

Why the replica broke: WAL, pg_basebackup and a silent hang

PostgreSQL replication streams the write-ahead log (WAL) from the primary to the replica. The primary keeps only a limited amount of WAL; if a replica falls too far behind, the segments it needs are recycled and it can never catch up by streaming. The standard safety net is WAL archiving: the primary copies every finished segment somewhere else, so a replica or a restore can fetch old ones. GitLab.com was not using WAL archiving, so the only fix was to wipe the replica's data directory and copy the whole database again with pg_basebackup.

That is where the night went wrong. pg_basebackup complained about max_wal_senders, which the engineer raised from 3 to 32. PostgreSQL then refused to restart because of too many semaphores: max_connections was set to 8,000, a value used for almost a year, and it was lowered to 2,000. Then pg_basebackup just sat there with no output.

Per the postmortem, pg_basebackup "will sit and wait silently" for the primary, up to 10 minutes according to another production engineer. That was not in the runbooks and not clearly in the docs. The engineer, who had said earlier that he would sign off around 23:00 local time, thought, in the live doc's words, "that perhaps pg_basebackup is being super pedantic about there being an empty data directory", and decided to remove the directory. In the terminal he was typing into, he was on db1.

The two hostnames, db1.cluster.gitlab.com and db2.cluster.gitlab.com, differ by one character. The postmortem's quote: "Unfortunately this process was executed on the primary instead. The engineer terminated the process a second or two after noticing their mistake, but at this point around 300 GB of data had already been removed."

Five PostgreSQL backups, none working

This is the part that made the incident famous. From the postmortem:

# Backup What they found
1 pg_dump to S3, every 24 hours "The S3 bucket was empty." The cron job ran on an app server with no PostgreSQL data directory, so the Omnibus package picked PostgreSQL 9.2 binaries for a 9.6 database, and pg_dump failed. The failure emails were rejected because DMARC was not enabled for cron mail
2 Azure disk snapshots Enabled for the NFS servers, not for the database servers: "we assumed our other backup procedures were sufficient"
3 The replica Its data directory had been wiped on purpose an hour earlier, for the re-sync
4 Daily LVM snapshot Almost 24 hours old, and the staging sync strips all webhooks from the copy
5 Manual LVM snapshot from ~17:20 Taken for the load-balancer test. About six hours old. This is the one they restored

"This means we were never aware of the backups failing, until it was too late." The next-day post summed it up in the sentence Hacker News quoted back in the 1,162-point live-report thread: "So in other words, out of five backup/replication techniques deployed none are working reliably or set up in the first place. We ended up restoring a six-hour-old backup."

Restoring meant copying the staging data directory back to production over Azure classic, non-premium network disks, throttled to about 60 Mbps. That took around 18 hours. Database sequences were incremented by 100,000 after the restore.

Blast radius and the public recovery

Everything written between about 17:20 and 23:25 UTC was lost: roughly 5,000 projects, 5,000 comments and 700 new user accounts (the live doc counted 5,037 projects, about 4,979 comments and 707 users). Git repositories and wikis live outside the database and were not affected; self-managed GitLab installations were not affected at all.

GitLab ran the recovery in public. The notes were a public Google Doc, updated live, and the restore was streamed on YouTube, with a peak of about 5,000 viewers; per the postmortem it was "the #2 live stream on YouTube for several hours". The live doc also has one of the most human lines in any incident record: the engineer "says it's best for him not to run anything with sudo any more today", handing the restore to a colleague.

The live GitLab.com incident doc (Wayback copy, Feb 2, 2017): the 22:00 replication-lag entries, the 23:00 removal on db1 instead of db2,

Who was to blame? The 5 Whys

Not the engineer. The postmortem keeps him anonymous and says GitLab "will redact names in future cases". The live doc had initials in it because he added his own. Its five-whys analysis ends on process: "Why was the backup procedure not tested on a regular basis? - Because there was no ownership, as a result nobody was responsible for testing this procedure."

My git blame for this one:

  • Two hostnames one character apart, in terminals that looked identical.
  • Five backup systems nobody had ever restored from. A backup that has never been restored is a hope.
  • Silent failures. A cron job whose errors go to email that bounces is the same as no cron job.
  • No owner for data durability. Each backup belonged to someone's setup work, and nobody's job was to check them all.

The postmortem also rejects the obvious fix: "one could alias rm to something safer but in doing so would only protect themselves against accidentally running rm -rf /important-data". Instead of preventing engineers from running commands, it makes the host obvious and the backups real. As ams6110 wrote on HN: "Good lesson on the risks of working on a live production system late at night when you're tired and/or frustrated."

How to avoid an rm -rf disaster on your own database

GitLab's fix list, with issue numbers, is still a good checklist: a per-host/environment shell prompt (#1094), Prometheus monitoring for backups (#1095), sane max_connections (#1096), point-in-time recovery with WAL archiving (#1097), hourly LVM snapshots (#1098, already running by Feb 10), Azure snapshots for database servers (#1099), automated restore testing (#1102), and an owner for data durability (#1163).

Two of them are small enough to do this week. A production prompt that cannot be mistaken for anything else (illustrative bash, put it in the production hosts' shell profile):

# red background, the word PRODUCTION and the full hostname
PS1='\[\e[41;97m\] PRODUCTION \[\e[0m\] \u@\H:\w\$ '
Enter fullscreen mode Exit fullscreen mode

And WAL archiving, so a lagging replica or a restore can fetch old segments instead of forcing a full re-copy (a simplified sketch of postgresql.conf; the archive location is yours to choose):

wal_level = replica
archive_mode = on
archive_command = 'test ! -f /mnt/wal_archive/%f && cp %p /mnt/wal_archive/%f'
Enter fullscreen mode Exit fullscreen mode

The rest is one habit: restore a backup, on purpose, on a schedule, and alert when it fails. If the restore has never been run, you do not know whether you have a backup.

Verdict: SHIP IT

I stamped GitLab's database incident SHIP IT, for the response. The outage itself was a stack of ordinary failures. What GitLab did next is why people still read it: the notes were public while it was happening, the recovery was streamed, the postmortem blamed the process instead of the person, the CEO apologised by name ("I apologize personally, as GitLab's CEO"), and the fixes shipped with issue numbers you could follow. Monday action: restore a backup.

FAQ

What happened in the GitLab database incident?
On January 31, 2017, an engineer deleted the production PostgreSQL data directory while trying to re-sync a replica. None of the five backup methods worked, so GitLab restored a six-hour-old snapshot and was down for about 18 hours.

How much data did GitLab lose?
About six hours of database writes: roughly 5,000 projects, 5,000 comments and 700 new user accounts. Git repositories and wikis were not affected.

Why did GitLab's pg_dump backups fail?
The cron job ran on a server that picked PostgreSQL 9.2 binaries for a 9.6 database, so pg_dump failed, and the failure emails were rejected because DMARC was not enabled.

Can you undo rm -rf?
Not with rm itself; there is no trash. You restore from a backup, which is why the backup has to be tested.

Sources


This article expands on an episode of **The Daily Diff, a five-minute daily video on what shipped and what broke in tech.
Watch the episode · Subscribe on YouTube · the written diff lands in your inbox every morning at thedailydiff.dev.

Top comments (1)