DEV Community

Bogdan
Bogdan

Posted on

Your AI agent re-sends the email on retry: an outbox for side effects

Most agent frameworks model a step as "do some work, then call the tools."
So the tool call — send the email, call the webhook, charge the card — happens
inside the run, somewhere in the middle of your graph.

That works until it doesn't. The moment you add a retry, a crash-resume, or a
deterministic replay, the same step runs again, and the side effect fires again.
State you can version and reproduce; the world you can't.

This post is about the outbox I shipped in
reactifact 0.14.0 — still the design in
the current 0.15.x line. The produce records what should happen as an
artifact in the context (committed atomically with everything else); the runtime
delivers it after the commit, once per stable id. Replay reconstructs the
answer without re-sending anything.

TL;DR — Don't do I/O inside an agent step. Record the intent as state
(committed atomically with everything else) and let the runtime deliver it
once, after the commit. Retry, resume and replay then rebuild the decision
without re-sending.

Runnable in about a minute, no API key:

pip install reactifact
python -m examples.outbox.main   # seven cases, each one asserted
Enter fullscreen mode Exit fullscreen mode

Code: github.com/bzdvdn/reactifact

If you know the graph frameworks, the difference is where the effect lives:

Typical agent step reactifact
The step does calls the tool mid-run writes the intent as state
On retry / resume fires the side effect again re-derives the same stable id → no second send
On replay re-executes → re-sends reconstructs the record → never re-sends
On parallel/merge two commits, two sends one intent, one delivery

The problem, concretely

Here's the shape that causes trouble. A produce reacts to an order and sends a
confirmation email:

async def process_order(order):
    email.send(order.customer, "Your order is confirmed")   # side effect, mid-run
    return Receipt(order_id=order.id)
Enter fullscreen mode Exit fullscreen mode

Now consider the three things every long-lived agent eventually needs:

  • Retry / resume. A crashed run resumes from its last checkpoint and re-runs the step. The email goes out twice.
  • Parallel producers. Two produces in the same generation both look at the pre-commit snapshot, both see "no receipt yet", both send.
  • Replay. You want to answer "why did the agent send this?" — but a replay that re-executes the step re-sends the email, so it isn't a replay at all.

You can try to guard it. A create_once(...)-style check helps for state, but the
guard resolves at commit time, and two producers in one generation share the
same pre-commit snapshot — so both pass the guard and both send. The check is in
the wrong place: it's protecting the write, while the side effect happens
before the write.

Intent is state; delivery is not

The outbox inverts the order:

  1. The produce records an intent — an artifact called PendingAction — and does no I/O.
  2. The runtime compiles the produce's effects (the intent and anything else) into one atomic patch and commits it.
  3. After the commit, the runtime hands committed-but-undelivered intents to a dispatcher, once per stable id, and records the outcome as state.

A produce records a PendingAction (no I/O); the runtime compiles and commits it; only after the commit does a dispatcher call the world once per stable id. Replay, retry and merge rebuild state from the commit chain and never re-send.

In reactifact that reads like this:

from reactifact import PendingAction, ProduceCall, produce


@produce(Receipt, also_creates=[PendingAction])
async def process_order(call: ProduceCall) -> None:
    order = call.trigger

    # The outbound side effect: recorded, not performed. The key is derived
    # from the order id, so a re-run or a merged branch reuses the same intent.
    call.effects.act(
        "notify",
        key=f"notify:{order.data.id}",
        payload={"to": order.data.customer, "order": order.data.id},
    )
    call.effects.upsert(
        Receipt(order_id=order.data.id, text="confirmed"),
        id=f"receipt:{order.data.id}",
    )
    return None   # nothing is applied until the runtime compiles the effects
Enter fullscreen mode Exit fullscreen mode

effects.act(...) creates a PendingAction under the stable id
action:{key} and returns None if one already exists (an idempotent re-run).
Note what the produce does not do: it never touches the network. It only states
a change.

Delivery is a separate, injected step:

async def dispatch(context, action):
    # a real one calls your email/webhook API here
    if action.data.idempotency_key in sent:
        return
    sent.append(action.data.idempotency_key)


runtime = Runtime(ctx, agents=[Notifier()], dispatcher=dispatch)
await runtime.arun()
Enter fullscreen mode Exit fullscreen mode

The runtime drains the outbox after each generation's commit, marking each action
dispatched — or failed (and re-raising) if the dispatcher throws. A failing
dispatch is state, not a lost effect.

What the split buys you

  • Replay reconstructs, never re-sends. A reactifact replay walks the commit chain and rebuilds the Context without running the runtime. The recorded notification is read back as state. Replay answers "why did it send that?" without sending it again.
  • Retries don't duplicate. A resumed run re-derives the same stable id, so effects.act returns None and no second intent exists.
  • Merged branches converge. Two forks that independently reach the same action share the id and merge to one intent. The dispatch bookkeeping (status/timestamp/error) is excluded from the three-way merge signature, and the dispatched side wins regardless of which branch is the merge target — so a merge can never resurrect an already-sent action. Divergent payloads under the same id are an explicit merge conflict, not a silent last-write-wins.
  • Same-generation duplicates collapse. Two producers that both record the intent in one generation produce one artifact and one delivery, because the dedupe happens at drain, after the commit.

You get the operational story you'd want from a queue — but the "queue" is just
versioned state you already have.

The honest limits

This is deliberately not a message broker, and the limitations are worth stating
plainly:

  • It's at-least-once, not exactly-once. A crash between the real send and the dispatched commit re-runs the delivery. The framework cannot make the I/O exactly-once — so the intent carries an idempotency_key you pass to the external system, which dedupes on its side. That is the same contract every real outbox relies on.
  • There is no background relay. The next arun() or an explicit flush_pending_actions() drains the outbox; the framework doesn't run a poller for you. Retry and backoff policy stay with the application (wrap your dispatcher), because that's a product decision, not a framework reflex.
  • It's single-process. The outbox makes state replay-safe; it does not turn reactifact into a distributed task queue.

That's the point of the split: state is versioned and reproducible, the world
isn't, and the boundary between them should be something you declare rather
than something buried in the middle of a function.

Run it offline

The whole thing ships as an executable, no-API-key demo. examples/outbox walks
seven cases and asserts each one:

.venv/bin/python -m examples.outbox.main
Enter fullscreen mode Exit fullscreen mode
1. commit -> dispatch        sent=['notify:42'] status=dispatched
2. re-derivation             first run sent 1; the re-run sent 0
3. same generation           producers=2 intents=1 sent=1
4. two merged branches       branches=2 intents-after-merge=1 sent=1
5. replay                    sent before replay=1 after replay=1
6. failure -> retry          raised 'smtp temporarily unavailable'; retry sent once
7. app-owned retry           transient outage absorbed by the wrapper
Enter fullscreen mode Exit fullscreen mode

It shipped in 0.14.0, alongside correlated structured logging and a hard
per-turn deadline with graceful shutdown (the current release is 0.15.x):

That's half the story: the outbox keeps an external effect from firing twice.
The other half — making the state behind an answer inspectable and
reproducible, so you can hash it, replay it and audit the trace — is the
next post.

If you're building agents that touch the outside world, I'd genuinely like to
know: where do you draw the line between a replayable computation and an
external effect?
Intents-as-state is one answer — I'm curious about others.

Top comments (0)