Hello, I'm Maneshwar, and I'm building LiveReview — a blast-radius aware AI code review built for your business-critical systems. Star us to help devs discover the project, give it a try, and share your feedback to help improve the product.
When you are first building an application, a transaction is the cheapest guarantee you will ever get.
A customer places an order. You charge their card, reserve the inventory, write a ledger entry for accounting.
You wrap the whole thing in BEGIN and COMMIT, and if any one of those writes falls over, the database quietly rewinds the other two.
It is as if the order never happened. You probably wrote that code in four minutes and never thought about it again.
What you were leaning on there are two of the ACID guarantees, doing unpaid labour on your behalf.
Atomicity. All three writes land, or none of them do. There is no universe where the card gets charged but the stock never gets reserved.
Isolation. While the transaction is in flight, nobody else can see the half-finished middle. Another request checking the customer's balance will not catch a charge for an order that has not finished processing.
You wrote the SQL. The database did the genuinely hard part, for free, and never asked for credit.
Then you split the database, and the guarantee quietly leaves
Traffic grows. Writes grow. That one box starts breathing heavily.
So you do what everyone does. You shard, or you break the monolith into services and give each one its own database.
The details vary. The outcome is identical: your data now lives on several independent machines that have never been introduced to each other.
That one payment flow is now three separate operations against three separate databases.
Card charged in the payments database. Stock reserved in the inventory database. Ledger entry recorded in accounting.
There is no BEGIN big enough to cover all three, because none of them knows the others exist.
So when the card charge commits and the inventory reservation then fails because the last unit just sold, nothing rolls back.
That charge is already durable, on a different machine, in a database that has no idea an order died somewhere else.
At a few hundred orders a day, you will see this maybe once and fix it by hand. At a few thousand transactions a second across distributed infrastructure, partial failure stops being an incident and starts being a Tuesday.
This whole shape of problem has a name: the distributed transaction. One logical operation spanning multiple independent databases or services, where every step needs to either succeed together or get cleaned up when something goes sideways.
The literature offers you two answers. The industry has very firmly picked one of them, and knowing why will save you a genuinely painful quarter.
The textbook answer: two-phase commit
Two-phase commit is the classic academic solution, formalised decades ago in the X/Open XA specification.
You introduce a new component, a coordinator, whose entire job in life is to make every participant agree on the outcome before anybody makes anything permanent.
It runs in two phases, which is where the name comes from.
Phase one, prepare. The coordinator asks every participant: can you commit this? Each database does the real work, durably records the change so a crash cannot lose it, locks the affected rows so nothing else can touch them, and then answers yes or no.
Phase two, commit. If anyone voted no, the coordinator tells everyone to abort and release their locks. If everyone voted yes, it tells everyone to commit, the locks come off, and you are done.
What you get out of this is the same strong consistency you had with one database. Every participant agrees before anything is finalised, so there is no visible window where the system is half-done.
On a whiteboard, this is exactly what you want. It is beautiful. It is also why almost nobody runs it across services.
Blocking is the whole problem
2PC is a blocking protocol, and blocking in a distributed system means you have made your availability the product of everybody else's uptime.
Picture the coordinator collecting all three yes votes. Then, in the gap between collecting them and broadcasting the decision, it dies.
Now look at the participants.
They cannot commit on their own, because the coordinator might have been about to say abort.
They cannot abort on their own either, because the coordinator might have been about to say commit, and one of their peers might have already gone through with it.
So they wait. Holding locks. Indefinitely.
And every other transaction in your system that wants to touch any of those rows is now queued behind a decision that nobody alive can make.
Crashes are not even the interesting failure. The boring one is worse.
A single slow participant stalls the entire transaction. If the ledger service takes ten seconds to answer the prepare message, the payment service and the inventory service both sit there holding locks for ten seconds, doing nothing productive whatsoever.
Your whole system now moves at the speed of its slowest participant, which in a microservice estate is whichever team deployed most recently.
Add a network partition and it gets worse still, because the coordinator cannot distinguish "the message did not arrive" from "the message arrived and the reply got lost". There is no safe default to pick.
Pat Helland's Life beyond Distributed Transactions made this argument in 2007 and it has aged extremely well. Distributed transactions across autonomous services do not work at internet scale, so build systems that do not need them.
Where 2PC actually does live
To be fair to the protocol, 2PC is alive and shipping. It is just not shipping where you would write it yourself.
It lives inside distributed databases. Google Spanner and YugabyteDB both run two-phase commit under the hood to coordinate writes across shards.
The difference is coupling. In there, the coordinator and the participants are components of one system, with one deployment, one failure model, one team, and a consensus protocol like Paxos or Raft underneath to make the coordinator itself fault-tolerant.
Crucially, the database is absorbing that complexity so that you, the caller, never see it.
The moment you try to rebuild that arrangement across independent services with different release trains, different on-call rotations and different opinions about timeouts, every assumption that makes it work in Spanner evaporates.
Sagas: give up atomicity, keep your weekends
When companies genuinely need to coordinate work across services, the pattern they reach for is the saga. Uber, Netflix, Amazon and DoorDash all run this in production.
Sagas start from a deliberately humbler assumption.
You do not actually need all-or-nothing atomicity spanning five services. You need a way to reliably end up in a consistent state, even when a step blows up along the way.
So instead of one big locked distributed transaction, you break the work into a chain of independent local transactions. Each service does its piece and commits to its own database, on its own terms, immediately.
When something fails halfway down the chain, there is no rollback available, because the earlier steps are already committed somewhere else. Instead you run a compensating action.
That is a business-level undo. A refund instead of a rollback. A cancellation instead of an abort.
What you trade away is strong consistency. What you get is eventual consistency.
The system can genuinely be inconsistent for a few seconds. The customer might see a charge land before they see the refund arrive.
But it always converges, and nothing is blocked while it converges. Every other transaction in the system keeps flowing the entire time, because no saga step ever holds a lock that a stranger on another machine is responsible for releasing.
That is the actual trade. Not "correct versus incorrect", but "briefly visible mess" versus "system-wide stall".
Choreography or orchestration
Something has to notice the failure and kick off the compensations. Who that something is turns out to be the design decision that matters most.
Choreography is the decentralised option. It is publish and subscribe: each service broadcasts an event when it finishes, and whoever cares picks it up.
Payments charges the card and publishes CardCharged. Inventory is listening, reserves the stock, publishes StockReserved. Ledger picks that up and records the entry. On failure, a service publishes a failure event and upstream services run their own compensations.
Orchestration is the centralised option. A dedicated orchestrator drives the flow, calling one service at a time and waiting for confirmation before moving on.
Choreography is lovely at three steps. It is miserable at eight.
Once a dozen services are publishing and reacting to each other, "where is order 1841 right now?" becomes an archaeology project. Which step failed? Which compensations already ran? Did the refund actually go through, or is it still retrying?
Nobody owns the answer, so you go and read twelve services' logs to reconstruct it.
Orchestration costs you a central component, and buys you the one thing choreography cannot give you: a single place that knows the state of every in-flight transaction.
Temporal, built by the people behind Uber's Cadence workflow engine, and AWS Step Functions both exist specifically to be that component.
And here is the part that matters most, the reason the orchestrator is not just 2PC wearing a new hat.
When a saga orchestrator crashes, nothing is holding locks.
It is durable. It restarts, reads its own state back from its own database, and resumes from exactly where it stopped. Meanwhile the rest of your system never noticed, because no rows were ever frozen waiting on it.
A dead 2PC coordinator freezes your data. A dead orchestrator just delays one workflow.
Compensating actions are where the honesty comes in
"Just undo the previous step" sounds tidy. In practice, it is where sagas get their reputation.
The card charge committed, the inventory reservation failed, so you refund. Fine. Except that refund is not the invisible cleanup a database rollback gives you.
The customer sees a charge appear on their card. Then, seconds later, a refund. Their bank probably sent them a push notification for each one. It is correct, and it still generates a support ticket.
Some actions are worse than visible. Some are simply not undoable.
You cannot unsend an email. You can send a follow-up saying "ignore that last one", which is a different thing pretending to be the same thing.
If you fired a webhook at a third party, you can send a cancellation, but you cannot make them process it in time, or at all.
Every step in your saga needs a compensating action defined up front, and you should be clear-eyed that some of them are inherently imperfect.
Then there is the recursive problem: compensations fail too.
What happens when the refund API is down at exactly the moment you need to refund? Now you need retries on your failure handling. And the instant you have retries, you need idempotency, otherwise a flaky network refunds the customer twice and you have invented a new bug that is much harder to explain in a postmortem.
The trick is to derive the idempotency key from the thing you are undoing, never from the attempt:
def compensate_charge(order_id, amount_cents):
# keyed on the order, not on this attempt, so ten retries of this
# compensation collapse into exactly one refund on the customer's card
key = f"refund:order:{order_id}"
return stripe.Refund.create(
payment_intent=lookup_payment_intent(order_id),
amount=amount_cents,
idempotency_key=key,
)
Generate a fresh UUID per attempt and you have an idempotency key that guarantees nothing.
The uncomfortable summary is that your unhappy path needs the same reliability engineering as your happy path. Most teams budget for one of those.
The dual write problem, quietly waiting behind all of this
Even with perfect compensation logic, there is one more failure mode that catches teams off guard, and it is the one I would bet money is already lurking in your codebase.
When the payment service finishes charging the card, it has to do two things. Save the result to its own database, and publish an event so the next step knows to run.
Those are two writes, to two different systems, with no shared transaction between them.
If the database write succeeds and the event publish fails, the next step never fires and the saga stalls silently. No error, no compensation, just an order sitting in charged forever until a customer emails you.
If the event publishes and the database write fails, downstream services are now reacting enthusiastically to something that never happened.
The fix is the transactional outbox, and it is one of those patterns that feels like cheating the first time you see it.
Stop treating the event as a separate write. Put it in the same database, in the same transaction, as the data.
BEGIN;
UPDATE payments
SET status = 'charged'
WHERE order_id = 1841;
INSERT INTO outbox (aggregate_id, event_type, payload)
VALUES (1841, 'CardCharged', '{"order_id":1841,"amount_cents":4999}');
COMMIT;
One local transaction. Both rows commit, or neither does. The dual write is gone, because there is no longer a second write to fail.
A separate background relay then reads that outbox table and publishes to the broker. It can tail the database's own transaction log via change data capture, which is what Debezium does, or it can just poll the table on an interval, which is unglamorous and works fine at most scales.
flowchart TD
A[Card charged successfully] --> B[BEGIN local transaction]
B --> C[Write the payment row]
C --> D[Write the event into the outbox table]
D --> E[COMMIT: both rows, or neither]
E --> F{How does the relay<br/>pick it up?}
F -->|CDC, tail the WAL| G[Publish to the message broker]
F -->|or just poll the table| G
G --> H[Next saga step runs]
classDef decision fill:#f4d35e,stroke:#b8991f,color:#1a1a1a
classDef start fill:#e9ecef,stroke:#6c757d,color:#1a1a1a
classDef db fill:#5ee6c8,stroke:#1f9c86,color:#1a1a1a
classDef out fill:#6ea8ff,stroke:#2f5fbf,color:#1a1a1a
class A start
class F decision
class B,C,D,E db
class G,H out
Note that this gives you at-least-once delivery, not exactly-once. The relay can crash after publishing and before marking the row as sent, so the event goes out twice.
Which brings us back to idempotency. Every consumer in your saga needs to handle seeing the same event twice, because it will.
So what should you actually build?
Before either of these patterns, ask the question that people skip because it feels like a cop-out.
Do you need a distributed transaction at all?
If you can draw your service boundaries so that data which changes together lives in the same database, do that. Move the inventory and ledger tables next to the payments table if they are always updated together.
A local transaction is simpler, faster and more reliable than any distributed thing you can build on top of it. This is always the best answer when you can get away with it, and it is enormously easier to get right up front than to retrofit two years and forty services later.
flowchart TD
A[One operation spans several services] --> B{Can that data live<br/>in a single database?}
B -->|Yes| C[Just use a local transaction]
B -->|No| D{Is eventual consistency<br/>acceptable here?}
D -->|No| E[Put it in a distributed SQL database<br/>Spanner or YugabyteDB]
D -->|Yes| F{Many steps, branching,<br/>or tricky compensations?}
F -->|No| G[Saga with choreography]
F -->|Yes| H[Saga with orchestration]
H --> I[Temporal or AWS Step Functions]
G --> J[Transactional outbox<br/>+ idempotent consumers]
H --> J
classDef decision fill:#f4d35e,stroke:#b8991f,color:#1a1a1a
classDef start fill:#e9ecef,stroke:#6c757d,color:#1a1a1a
classDef good fill:#5ee6c8,stroke:#1f9c86,color:#1a1a1a
classDef alt fill:#9d8cff,stroke:#5b4bcc,color:#1a1a1a
classDef tool fill:#6ea8ff,stroke:#2f5fbf,color:#1a1a1a
class A start
class B,D,F decision
class C,H good
class G,E alt
class I,J tool
If you truly cannot keep it in one database, you are writing a saga. That part is not really up for debate anymore.
The remaining question is which flavour.
Choreography if the flow is short, the services are genuinely independent, and nobody needs a central view of where a transaction stands. An order placed event triggering an email and a push notification is a perfect fit, and wiring up an orchestrator for it would be theatre.
Orchestration for anything with branching logic, more than a handful of steps, or compensation logic you would rather define in one file than scatter across six repos. Most teams end up here eventually, and the tooling has got good enough that starting here is no longer expensive.
And if eventual consistency is genuinely unacceptable for some specific slice of your data, put that slice in a distributed database that handles strong consistency internally. That is a completely different proposition from hand-rolling 2PC across services you do not control.
The shape that keeps winning
Strip it all back and the production answer at basically every company operating at scale looks the same.
A saga, orchestrated by something durable. Idempotent operations at every step, so retries are always safe. A transactional outbox, so events are as reliable as the database writes they describe.
It means accepting that your system is briefly, visibly inconsistent sometimes.
That is not a compromise the industry stumbled into. It is one it chose, on purpose, after watching two-phase commit deadlock enough production databases to make the decision for everyone.
The database gave you atomicity for free. The moment you split it, the bill came due, and sagas are simply the cheapest way anyone has found to pay it.
Your team's attention is limited, and the deluge of AI-generated code is making it harder to keep production secure and reliable without slowing you down.
I'm building LiveReview, a blast-radius aware AI code review built for your business-critical systems.
Instead of presenting every diff with equal emphasis, LiveReview scores each change by blast radius — how far its impact reaches through your call graph — so you can focus attention where it actually matters.
Spend code review effort where business risk is highest — not spread evenly across every diff.
⭐ Star it on GitHub:
HexmosTech
/
LiveReview
Blast-Radius Aware AI Code Review for Business-Critical Systems
LiveReview: Blast-Radius Aware AI Code Review for Business-Critical Systems
LiveReview is an AI code reviewer that scores every hunk of a diff by blast radius: how far a change reaches through your call graph, how much persistent state it touches, and how well-tested it is. A 3-line change to a shared auth check can outrank a 300-line UI tweak. Your team's attention goes to the highest-risk code first, not spread evenly across every diff.
blast-radius-demo.mp4
LiveReview's Blast Radius & Review Priority scoring, live in the diff viewer.
Here's the goal:
- A 3-line fix in a function used by 40 other files, that also writes to a database, should score high.
- A 300-line UI change in one file, fully covered by…
Click below to try LiveReview with your codebase:












Top comments (0)