DEV Community

Rodolfo Henrique
Rodolfo Henrique

Posted on

The 3-2-1-1-0 backup rule. We had everything except the zero

The 3-2-1-1-0 backup rule. We had everything except the zero

Backup looks simple until the day you need it.

In simple terms, a backup means making a copy of your data so it can be recovered if the original is lost, corrupted, or deleted.

But there is an important difference between having a backup and being able to recover your data from it.

Imagine that every day an automated process backs up your database. At the end, the system shows a big green mark: the process finished without errors.

You might think: "everything is protected."

But what does that green mark actually mean?

Only that the process was able to run the steps it was supposed to run. By itself, it does not guarantee that the copy is intact, that the data is consistent, or that you can recover it when you really need to.

This is where the 3-2-1-1-0 rule comes in.

It is a strategy to reduce the risk of losing data even when something goes wrong. The numbers describe how copies should be kept and protected:

  • 3 copies of the data
  • 2 different types of media or storage
  • 1 copy kept off the primary environment
  • 1 isolated or immutable copy
  • 0 errors during the recovery test

The last number is the interesting one.

Zero is not another copy or another place to store backups. It is a requirement: when it is time to recover the data, the restore has to work.

Because a backup that has never been restored is, to a point, a hypothesis.

The scene below is the starting point for understanding that count, and especially the reasoning behind each number.

Backup protection strategies: 3 copies of the data, 2 different types of media or storage, 1 copy off the primary environment, 1 isolated or immutable copy, and 0 errors on the recovery test.

What fails when a number is missing

The rule only holds if each number covers a different failure. If two numbers break the same way, you do not have five controls. You have repetition.

Without the 3, production and the "backup" live in the same place. A retention mistake, a broad policy, or a compromised identity deletes the original and the copies together. Three names in one location are not three copies.

Without the 2, the copies use the same media or the same storage service. The failure repeats: corruption, quota, API, region. Two identical disks in the same account fail the same way.

Without the 1 off the primary environment, the incident in that environment takes the backup with it. Compromised account, single IdP, unavailable region. The copy has to sit where that incident cannot reach.

Without the 1 that is isolated or immutable, whoever reached production also reaches the backup delete. Isolation cuts the easy path. Immutability blocks the delete even when the path appears. Ransomware that hits production usually hits the backup repository next. Without that lock, the way back disappears with the original.

Without the 0, the job is green and the box will not open. The first four numbers say where the copy lives. The zero says whether it comes back.

That is failure engineering, not a poster. It is still incomplete until restore has been exercised.

The green job lies by omission

A successful backup job answers a narrow question: the writer persisted what was requested, to the configured destination, inside the window. It does not answer whether the incremental chain is intact. It does not answer whether the database was captured in a usable state (with a flush, a consistent snapshot, a native backup), or only as a disk in the middle of a write. It does not answer whether the KMS key for restore still exists, whether the identity doing the restore can read the vault, whether the destination engine accepts that version, or whether the time to return fits what the business calls acceptable.

Integrity means the backup is complete and coherent. Availability means you can reach it and apply it. Neither one is the job's exit code.

That is why the zero is not a quantity. It is not "one more vault". It is the acceptance criterion: restore on purpose, watch the data come back, measure how long it took, and only then call the design 3-2-1-1-0. Test the job. Test recoverability. Both. One without the other leaves a green check on a box that will not open.

The backup job is green. The restore failed. A copy is not recovery.

What a recovery test has to show

The zero does not ask for a lab. It asks for four answers, obtained on purpose, without panic:

  • did the data come back?
  • did it come back consistent, in a state the application will accept?
  • could someone on the team run the restore without inventing permissions, keys, or a destination on the spot?
  • does the time to return fit what the business can stand?

If any answer is "the job was green", the test has not happened yet. A dashboard measures the writer. Recovery measures the way back.

Does the backup work when you need it?

"Do we have a backup?"

That sounds like the right question. In practice, it says little. Most environments can answer yes.

The question that actually matters is this one:

When was the last time someone restored that backup on purpose, in a real environment, and confirmed that the system came back up?

A dashboard can show that the backup finished successfully. It can show that the copy exists, that storage is healthy, and that the last job completed without errors.

None of those signals answers whether you can recover the system when you actually need to.

If the only evidence that the backup works is a green dashboard, you know you have a copy. You might already have the 3, the 2, and both 1s.

But you are still missing the zero.

The zero is proof that the way back works.

Top comments (1)

Collapse
 
dhruv_malaviya profile image
Dhruv Malaviya •

"A backup that has never been restored is a hypothesis" is the line that should be on a wall, and the zero is the number everyone skips.

Worth mapping the other four onto a real setup. At Krova Cloud snapshots go to S3-compatible object storage separate from the bare-metal host the Cube runs on, encrypted with a key unique to each Cube — that's your offsite copy, and losing the host doesn't take the history with it. They're system-managed and can't be deleted by hand, which covers accidental tidying.

The part I'd push hardest is your zero: restore is one click, or a clone into a new Cube, so a recovery test costs nothing. An untested restore is the expensive kind.

How often do you actually run the restore?