DEV Community

Posting Dude
Posting Dude

Posted on

Your Backups Don't Exist Until You've Restored One: A Monthly Drill for a Solo SaaS

Ask a solo SaaS founder about backups and you usually get the same answer. "Yeah, the host does daily backups." Then you ask when they last restored one, and it gets quiet.

A backup you never restored is a guess. It might be fine. It might be an empty bucket.

That's not me being dramatic. In January 2017 GitLab.com lost about six hours of production database data after an engineer wiped the wrong Postgres data directory. They had several backup methods on paper. When they went looking, the nightly pg_dump uploads to S3 weren't there. The job had been running pg_dump 9.2 against a 9.6 database, so it errored out every time, and the failure emails were getting rejected by the receiving server, so nobody saw them. They recovered from a snapshot an engineer had taken by hand around six hours before for staging. One line in GitLab's postmortem stuck with me: the backup procedure wasn't tested regularly because nobody owned it.

If a company that size can have zero working backups for a while, a one-person SaaS definitely can. So here is the small drill I'd run once a month. It takes about an hour the first time and maybe 20 minutes after that.

First, write down what "a backup" means for you

Before any commands, answer three boring questions in a note:

  1. Where is the data? Usually Postgres (or MySQL), plus file uploads in S3 or similar, plus whatever lives only in Stripe. People forget the uploads bucket all the time.
  2. How much can I lose? If you restored last night's backup right now, could you live with losing everything since? For most early products a day is survivable. An hour is nicer. Be honest, not heroic.
  3. How long can I be down? "A few hours on a weekend" is a real answer. It decides whether a slow manual restore is ok.

That's your recovery point and recovery time, in plain words. Don't skip it, because it tells you whether the host's daily snapshot is actually enough.

The monthly restore drill

1. Get a real backup, not the live database

Use whatever your provider gives you (daily snapshot, point-in-time restore) or your own dump. If you run your own, the custom format is the one to use, because pg_restore can list it and restore pieces of it:

pg_dump -Fc "$DATABASE_URL" > app-$(date +%F).dump
Enter fullscreen mode Exit fullscreen mode

Check the pg_dump version matches your server's major version. pg_dump --version vs SELECT version();. This is literally the bug that ate GitLab's dumps.

2. Restore into a throwaway database, never over prod

Spin up a scratch database (a local Docker Postgres, a new branch, or a temporary instance from the provider's restore button) and load the backup there:

createdb -T template0 restore_test
pg_restore --no-owner -d restore_test app-2026-10-06.dump
Enter fullscreen mode Exit fullscreen mode

template0 gives you a truly empty database, and --no-owner stops it failing on roles that don't exist on your laptop. If you want to see what's inside before loading anything, pg_restore -l app-2026-10-06.dump prints the table of contents.

3. Check it like a suspicious customer

"It restored without errors" isn't the check. Run a few queries you already know the answers to:

SELECT count(*) FROM users;
SELECT max(created_at) FROM users;
SELECT count(*) FROM subscriptions WHERE status = 'active';
Enter fullscreen mode Exit fullscreen mode

Compare with prod. The max(created_at) one is the important one: it tells you how old the backup really is. Then open one real customer's record and make sure their stuff is all there, not just the row in users.

If you store uploads separately, pick three random file keys from the restored database and confirm those files exist in the bucket (or the bucket's backup).

4. Time it

Write down how long the whole restore took, start to finish. If it's three hours and you said you can be down for one, you've learned something important on a quiet Tuesday instead of during an outage.

5. Write the runbook while it's fresh

Five or six lines in your repo, docs/restore.md:

  • where the backups live and who can access them
  • the exact commands you just ran
  • which env vars / secrets you needed
  • how long it took
  • the date of the last successful drill

Future-you at 2am, half asleep and panicking, should be able to copy-paste it. That's the whole point.

6. Delete the scratch database

It has real customer data in it. Don't leave a copy on your laptop or in a forgotten instance.

The traps I'd watch for

Silent failures. A cron job that dumps the database and emails you when it fails is only as good as that email reaching you. Better: something that pings you when the backup doesn't succeed, like a dead man's switch (a heartbeat check that alerts if the job doesn't check in). Even better: the monthly drill, because it checks the thing you actually care about.

One copy in one account. If the backups live in the same cloud account as prod, one leaked key or one billing mess can take both. Keep at least one copy somewhere else, even a weekly dump to a second provider.

Replicas aren't backups. A replica copies your mistakes in real time. DELETE FROM users with no WHERE lands on the replica a second later. GitLab's standby wasn't usable either, because it got wiped while they were trying to fix the replication.

Restoring over production in a panic. Restore to a new database, check it, then switch the app over. Never pg_restore --clean into prod as your first move.

Why this pays off beyond disasters

Funnily enough, the drill also helps you sell. The first time a team buyer sends a security questionnaire, one of the questions is basically "do you have backups and have you tested them?" Being able to write "restored and verified on 6 Oct, takes about 40 minutes" is a lot stronger than "our host handles that". I wrote up what to put on a security page for your first B2B deals, and a tested restore is one of the cheapest lines you can honestly add to it.

And if a restore ever does happen for real, you'll need to tell people what's going on while you work. A simple status page and a habit of posting updates helps a lot there. I wrote about shipping a status page before your first real outage on Hashnode.

So, put a recurring calendar event on the first Tuesday of the month: "restore last night's backup". Do it once this week. Worst case you lose an hour. Best case you find out your bucket is empty while it still doesn't matter.

When did you last actually restore yours?

Top comments (0)