DEV Community

LakebaseGuru
LakebaseGuru

Posted on Originally published at Medium

I Fixed a Pricing Disaster With a Database Branch

Branch-based restores on Lakebase, told through a retail database.

Picture this. It is 1:47 AM on the Sunday before Thanksgiving week. Your team just pushed a pricing update to the Postgres database behind the e-commerce platform. By 2:15 AM the on-call engineer sees it: a bad join in the migration script overwrote regular prices with clearance prices across 40,000 SKUs. Every minute, orders keep flowing through at the wrong prices.

You have two bad options. Keep selling and bleed margin on the biggest week of the year, or take the storefront offline while the database team runs the classic recovery sequence: provision storage, copy the backup, replay WAL, wait. On a database this size, the waiting is measured in hours.

There is now a third option. Branch-based restores on Databricks Lakebase went GA on October 1, and they change the shape of this night entirely. Instead of a restore procedure, you open a branch at 1:46 AM, one minute before the deploy. Seconds later it exists, with its own compute and its own connection string. You verify the prices are correct, repoint the application, and you are back. The corrupted branch stays right where it is for forensics. The downtime is minutes, and most of those minutes are spent verifying, not waiting.

Why a restore became a branch

The reason this works is architectural. Lakebase separates Postgres compute from storage, and compute ships write-ahead log records to a durable layer that keeps the full history. A restore request maps your timestamp to the corresponding log sequence number, creates a branch pointing at that moment, and attaches compute. Nothing is copied. Copy-on-write means the new branch shares all underlying data through pointers; only new writes take new storage.

The operational detail that matters: the original branch is never touched. It keeps serving whatever is still connected to it while the restored branch comes up alongside. Recovery and investigation stop being the same fraught operation. You bring the past online next to the present instead of replacing the present with the past.

Databricks reports recovering a 100 TB database in seconds this way. Treat that as a vendor number, but the order of magnitude is believable for a simple reason: no data moves, so recovery time stops scaling with database size.

The same operation, four more retail mornings

Once a restore is a branch, it stops belonging to disaster recovery and starts belonging to the weekly routine. Stay in retail for a minute:

The replenishment engine misfires. Your demand forecasting model pushes a bad set of purchase orders overnight. Before the next model run, branch the database. Now you can run the old and new forecasts side by side against identical data instead of arguing about what changed.

Finance needs last Tuesday. An auditor asks what inventory levels looked like at close of business Tuesday. Branch Tuesday, run the queries, drop the branch. No ticket to the DBA team, no snapshot lifecycle to manage.

The AI agent goes sideways. You have an agent adjusting markdowns autonomously, which is exactly the kind of workload that makes experienced operators nervous. Snapshot the branch before the agent runs. If its pricing logic goes off the rails, the undo is a pointer flip, not a recovery project. Agents that write to production data need cheap instant undo the way code agents need git revert. Without it, every agent action is a one-way door.

The migration you were afraid to run. Schema change on the orders table during business hours. Branch first, run the migration against the branch, validate, then decide. The fear premium on routine database work drops toward zero.

The honest caveats

A few things to design around. The history window defaults to seven days and is configurable from two to thirty, so "any point in time" means any point inside your window. A restore covers every Postgres database on the branch, not a single one, and your application must be reconfigured to the restored branch's connection string. It is a new root branch, not an in-place rewind. Keep the window matched to your recovery objectives and treat the connection switch as part of the runbook.

The pattern underneath

First branching gave every pull request its own database. Now restores give every moment in history its own branch. The database is converging on the semantics developers already expect from git: cheap copies, instant reverts, history you can check out. In retail, where the database sits directly under revenue, that convergence is worth more than anywhere else. The teams that internalize it first will ship faster with less fear, because the cost of being wrong keeps dropping toward zero.

Further reading

Top comments (0)