Almost everyone has backups. Far fewer people have ever restored one.
That gap is where backup plans fail. The job ran every night, the dashboard showed green, and then the day came when someone needed the data and the files were empty, corrupted, or missing the one folder that mattered.
This post covers the 3-2-1 rule, a few ways backups quietly fail, and a small script you can adapt to check that your restores actually work.
The 3-2-1 rule
It's an old rule because it holds up:
- 3 copies of your data (the live copy plus two backups)
- 2 different types of storage (for example, a local disc and the cloud)
- 1 copy off-site, so a fire, theft, or flood can't take everything
A common modern addition is that at least one copy should be immutable or offline. Ransomware can encrypt anything it can reach, including a backup drive mounted on the same network.
How backups fail without anyone noticing
- The job runs but backs up the wrong thing. A moved folder or a new drive that was never added to the selection.
- The backup is corrupt. Silent errors on disk or in transit.
- Credentials expired. The cloud target rejects uploads and the alert goes to an unused inbox.
- Ransomware reaches the backup. It's writable from the same machine, so it gets encrypted too.
- Nobody knows the restore procedure. The data exists, but the one person who knew how to recover it left.
- Recovery is too slow. A full restore takes three days and the business can't wait that long.
Most of these only show up when you try a restore.
A simple restore check
The idea: restore the latest backup into a scratch folder, pick a few random files from the live data, and compare checksums against the restored copies.
Adjust the paths and the restore command for your tool. The example assumes a tool like restic that restores files under their original path inside the target folder.
#!/usr/bin/env bash
set -euo pipefail
SOURCE_DIR="/srv/company-data"
RESTORE_DIR="/tmp/restore-test"
SAMPLE_SIZE=5
rm -rf "$RESTORE_DIR"
mkdir -p "$RESTORE_DIR"
# 1. Restore the latest backup into a scratch folder.
# Replace this line with your backup tool's restore command.
# restic restore latest --target "$RESTORE_DIR"
# 2. Pick random files that haven't changed in the last day,
# so recent edits don't cause false alarms.
mapfile -t FILES < <(find "$SOURCE_DIR" -type f -mtime +1 | shuf -n "$SAMPLE_SIZE")
if [ "${#FILES[@]}" -eq 0 ]; then
echo "FAIL: no sample files found in $SOURCE_DIR"
exit 1
fi
failures=0
for f in "${FILES[@]}"; do
restored="$RESTORE_DIR$f"
if [ ! -f "$restored" ]; then
echo "MISSING: $f"
failures=$((failures + 1))
continue
fi
live_sum=$(sha256sum "$f" | awk '{print $1}')
backup_sum=$(sha256sum "$restored" | awk '{print $1}')
if [ "$live_sum" != "$backup_sum" ]; then
echo "MISMATCH: $f"
failures=$((failures + 1))
else
echo "OK: $f"
fi
done
rm -rf "$RESTORE_DIR"
if [ "$failures" -gt 0 ]; then
echo "Restore check FAILED ($failures problem(s))"
exit 1
fi
echo "Restore check passed"
Run it from cron or a scheduled task, and send failures somewhere a human will actually see them, such as email, a chat channel, or a monitoring tool.
What this script doesn't prove
- It checks a sample, not every file.
- It doesn't test databases or application consistency.
- It doesn't measure how long a full restore takes.
- It doesn't cover bare-metal recovery (rebuilding a whole machine).
Treat it as a smoke test, and still do a full recovery drill on a schedule.
A short checklist
- [ ] At least one copy is off-site.
- [ ] At least one copy can't be modified from your everyday network or accounts.
- [ ] Backup credentials are separate from daily-use admin accounts.
- [ ] Failure alerts go to a monitored channel.
- [ ] A sample restore runs automatically.
- [ ] A full restore drill happens at least a couple of times a year.
- [ ] The recovery steps are written down where someone else can find them.
- [ ] You know how long recovery takes and how much data you could lose
Where this shows up outside engineering
Small teams without a dedicated ops person often assume "the cloud" is handling it. That's worth checking. Many SaaS tools have limited retention or no easy way to recover from accidental deletion, and "synced" isn't the same as "backed up".
For a non-technical take you can share with a manager or client, see Cloud Backup vs Local Backup: Which Is Better for Small Businesses? and Is Your Business Data Really Backed Up?. If you'd rather hand it off, Zia Networks manages backups for small businesses in New Mexico.
Takeaway
Don't trust a backup until you've restored from it. Automate a small check, alert a human on failure, and practice the full recovery before you need it.
What does your restore testing look like today? Let me know in the comments.
Before you publish
Cover image: add one, since posts with images tend to get more clicks. A simple diagram of "3 copies, 2 media, 1 off-site" works.
Set published: true only after you've previewed it.
Don't set a canonical URL. This is original content, so it shouldn't point to your blog.
Test the script on a throwaway folder before you tell readers it works. Make sure the restore command matches your tool.
Tags: dev.to allows 4. Backup, security, DevOps, and beginners are all active.
Tips
Engage in comments. Replying to the first few comments helps the post get seen.
Keep the promotion light. I kept it to one short paragraph near the end, since Dev.to readers dislike ads.
Check link behaviour. Inspect the published page to see whether your links are followed.
Add a series. A follow-up on offboarding access or patching would build your profile.
Top comments (0)