DEV Community

informat
informat

Posted on

You Backed Up the Database. You Did Not Back Up the Application.

The IT manager at a plant in Dongguan called me on a Wednesday afternoon. Six hours earlier one of her admins had run a bulk update against an inspection table with the wrong filter attached, and sixty thousand rows now carried a blank where a measurement should have been.

"I do not need you to restore the server," she said. "I need you to restore that one table from last night. Everything else — the other thirty applications, the production dashboard, the approval queues — has to keep running. Can you do that?"

I told her I would get back to her. It took four days, and most of the work was manual.

That is the phone call I think about whenever someone asks whether our platform has backups. We did. There was a cron job that ran at two every morning and wrote two database dumps to a folder, exactly as our own documentation describes. What we did not have was a way to give her the thing she actually asked for.

A backup is a claim about a system

Here is the uncomfortable arithmetic. When you dump the platform database, you have captured part of the state. Not the whole of it.

Our platform keeps identity and permissions in one library and business data in another — two dumps, two files, two things that have to stay consistent with each other. In larger deployments the business library is already numbered, because it shards. The uploaded files, the scanned certificates, the photos people attach to records, live somewhere else entirely: object storage in a bucket, plus a local file directory that some installations still rely on. And the search index is derived, so we tell ourselves it does not need a backup — which is true right up until you restore and somebody notices the search box has gone quiet.

So "we back up the database" is a sentence that describes roughly half of the stateful parts of the system. It happens to be the half that is easiest to describe, which is why it is the half everyone stops at.

The missing pieces are not exotic. They are just boring enough that nobody owns them. The attachments are the most concrete example: a row restored without its file is a record pointing at an empty preview. The database says the certificate exists. The storage layer says it does not. Both are telling the truth about their own copy, and the user sees a broken link.

In low-code, the schema is data too

In a hand-written application, schema and data are different artifacts. They change at different rates, they are reviewed differently, and you can reason about restoring one without the other. That separation is the thing low-code quietly removes.

In our platform, adding a field is a metadata change. It also causes a physical column to appear. A table's physical name is a compound of two system-generated identifiers — stable for the table's whole life, and completely meaningless to a human reading the database. A "table" in the product is not a table in the database; it is an entry in a definition, plus a physical structure, plus indexes.

The consequence is that backing up the application and backing up its data are not two projects. They are one project, and most teams only discover that when they try to separate them. You cannot meaningfully restore yesterday's rows into today's definition without deciding which definition wins — and if you restore the definition, you have also just rolled back every field, rule and form change made since, for every user, not just the one who made the mistake.

This is why "restore the backup" is such a dangerous sentence to say to a non-engineer. It sounds like undo. It is closer to a time machine that moves everything.

The request is never "restore the instance"

Our documented recovery procedure is honest and it is blunt: before restoring, ensure the database has no active connections. That is a technical way of writing "stop the platform." Afterwards you restart the services, and then you fix what the restore dragged along with it — the licence is bound to the machine, so it needs re-authorising; the home address, login URL, mobile login URL, resource address and API address all get reset to whatever the source environment used; the document preview service points at the old host.

A restore, in other words, carries a whole environment with it. It is not self-contained, and it is not scoped. It is an all-or-nothing operation on an instance that is currently hosting forty applications for a dozen departments, at least one of which is probably running a shift handover.

That was the real gap. The customer's request was never "restore the instance." It was "restore this table, this time window, without touching anything else." Scope is the requirement, and scope is the part nobody builds, because building it means having an answer to a question we had never asked: what is the smallest unit of state you can name, isolate and put back?

Row-level operations are possible — you can extract a single physical table and reinsert it. But the trap is downstream. Every table carries its own sequence for record identifiers, and a numbered field carries a second one. Restore the rows and forget the sequence and the platform will happily hand out an identifier that already exists. Nothing fails that day. Something fails three weeks later, when a duplicate record number appears in a report and nobody can explain where it came from.

A delayed, silent, data-level failure is the worst kind of bug a platform can ship. And it is the natural result of treating a restore as a file copy instead of a reconciliation.

Retention, where backup and compliance disagree

There is a second contradiction, and it is harder to design around than the first.

Backup windows are short because storage is not free. Our own reference script for object storage keeps thirty days of copies and deletes the rest. Meanwhile the audit log is supposed to be long-lived evidence, the kind of thing an auditor asks for by year. And then a customer asks us to erase a person's data, completely, now, because a regulation says so.

You can promise durable point-in-time recovery. You can promise complete erasure. You cannot honestly promise both from the same copy of the data, because a dump is by definition a record you agreed to keep. Every platform that says yes to both is choosing which of the two sentences it is lying about, usually without noticing that it made a choice.

We stopped pretending. The backup exists to restore service, it has a documented lifetime, and erasure has a documented process that includes ageing out of the backup window. That is a less impressive answer than "we delete everything instantly." It is also a true one, and enterprise buyers can tell the difference.

The uncomfortable conclusion

We had treated backup as an operations concern — a cron job, two dumps, a folder, an scp to another machine. It is not. It is a product surface, and it is one of the few surfaces a customer only ever sees on their worst day.

Low-code is a machine for creating state with a drag. A business analyst can conjure a table, a rule, and an automation before lunch. If the platform cannot also name, scope, and restore that state, then it has handed the customer a way to create work that nobody can undo — and it will only find out when the phone rings on a Wednesday.

Backups are cheap.

Restores are the product.

Top comments (0)