I used to believe in rollback. Not casually. Religiously.
Every deploy had the same prayer: “If it breaks, we’ll just revert.”
Then one Thursday at 6:47 p.m., we shipped a “small” change. Forty minutes later, checkout was failing. Not completely down — just enough to be terrifying. Orders were getting stuck in pending. Customers were retrying. Support was already typing in all caps.
So we hit revert. CI went green. Everyone exhaled.
Then we checked the database.
The migration had already run. A new column existed. The old code ignored it, but a background worker didn’t. It kept reading the new column, finding nulls, and doing the wrong thing. We “reverted” for six hours while manually fixing rows and refunding duplicate charges.
That’s when I stopped treating rollback like time travel.
Rollback is not undo. It’s a new deploy that hopes it can clean up after the last one. And usually, it only cleans up the code. It doesn’t clean up:
- Data changes. Migrations, backfills, deleted rows, renamed columns.
- External side effects. Emails sent, payments charged, webhooks fired, push notifications delivered.
- State sitting in queues, caches, third-party APIs, and mobile apps still running last month’s build.
- Human decisions. Someone already saw the feature. Someone already changed a config. Someone already promised it to a customer.
The lie isn’t that revert is impossible. The lie is that revert makes you safe.
It makes you feel safe enough to skip the boring stuff. No feature flag. No canary. No backward-compatible migration. No one asks, “What happens if we need to roll back after data has changed?” Because the answer is always the same: “We’ll just revert.”
But that religion has a tax. Teams ship bigger changes because they think there’s a net. They skip dry runs. They treat staging like a ceremony. They forget that the hard part of rollback isn’t the code — it’s the state.
Real safety isn’t a revert button. It’s designing for forward fixes. Expand-and-contract migrations. Feature flags that actually turn things off. Canaries that expose 1% before 100%. Runbooks someone has actually tested. And enough humility to say, “If this goes wrong, we might not be able to go back. So let’s make the blast radius small.”
I still keep rollback in my toolbox. But I don’t worship it anymore. It’s a tactic, not a strategy.
So the next time someone says, “We can always revert,” ask them one question:
Revert what, exactly?
Top comments (0)