DEV Community

Cover image for The Saga Pattern Was a Workaround
Mark Fussell
Mark Fussell

Posted on Originally published at Medium

The Saga Pattern Was a Workaround

We spent a decade hand-building compensation logic because our runtimes couldn't remember where they were. Now they can.

The saga pattern is older than most of the systems it runs in. Hector Garcia-Molina and Kenneth Salem published it in 1987 for long-lived database transactions. The microservices era rediscovered it around 2015, and I spent a good part of the following decade explaining it on stage: when a business transaction spans services and step four fails, you can't roll back a distributed system, so you run compensating actions (refund the charge, release the inventory, cancel the booking) in reverse order.

What is the saga pattern?

A saga is a sequence of local transactions in which every step has a defined compensating action. If a step fails, the saga runs the compensations for every step that already completed. It trades the atomicity you lost when you distributed the system for eventual consistency that you manage yourself.

What a saga really is

Sagas work. Payment, travel, and retail systems run on them today. But be honest about what a saga actually is: a hand-built answer to the question "what happens when the process dies halfway through?"

To answer it, you maintain four things:

  • Compensation logic for every step
  • A state machine that tracks which step you're on
  • Recovery code that reads that state after a crash
  • Idempotency so a retried compensation doesn't fire twice

Every team builds this machinery, slightly differently, with the same bugs.

What durable execution changes

We built sagas this way because the runtime gave us nothing better. That premise has expired.

Durable execution journals every step as it runs. A crash resumes from the exact step that failed, completed steps never re-execute, and compensation becomes ordinary workflow code: recorded, replayable, and visible, instead of tribal machinery around the edges. The state machine you used to maintain by hand is now the engine.

Sagas don't disappear. The business reality they encode (some actions can only be undone, not prevented) is permanent. What disappears is the plumbing. On a durable workflow runtime, a saga is a dozen lines:

def order_saga(ctx: wf.DaprWorkflowContext, order):
    compensations = []
    try:
        yield ctx.call_activity(reserve_inventory, input=order)
        compensations.append(release_inventory)

        yield ctx.call_activity(charge_payment, input=order)
        compensations.append(refund_payment)

        yield ctx.call_activity(book_shipping, input=order)
    except Exception:
        for undo in reversed(compensations):
            yield ctx.call_activity(undo, input=order)
        raise
Enter fullscreen mode Exit fullscreen mode

Try the steps, catch the failure, run the compensations, all of it journaled. One of my favourite demos is deleting two-thirds of a team's saga code and watching the remaining third become readable.

Why agents raise the stakes

Agents make all of this matter more. An agent's compensations guard real money and real customer data. Its steps are expensive LLM calls you'd rather not re-run. And its failure points multiply with its autonomy: an agent that books a flight, charges a card, and then fails to reserve the hotel needs to unwind cleanly, not leave a customer half-booked.

If sagas were worth teaching through a decade of microservices, their durable-execution successor is worth learning in the first year of agents.

Further reading:

Top comments (0)