DEV Community

Cover image for How I built a time-travel debugger for LLM agents on Burr, and what "replay the unchanged branch and prove nothing changed" actually takes.
Harish Kotra (he/him)
Harish Kotra (he/him)

Posted on AI-assisted

How I built a time-travel debugger for LLM agents on Burr, and what "replay the unchanged branch and prove nothing changed" actually takes.

Rewind: counterfactual replay for agent pipelines

Every agent trace tool I've used has the same ceiling. It shows me a waterfall of steps, I spot that step six produced something stupid, and then I'm stuck. To find out why — or whether it was the model, the retrieval, or the prompt — I have to re-run the whole pipeline and hope it lands in the same place. With a non-deterministic model it doesn't, so I can never separate my change from model jitter.

So I built the other thing: fork a completed run at any step, change exactly one input, replay only the downstream steps, and diff the two trajectories. Rewind does that, and — this is the part that took real care — when you replay a branch you didn't change, every output hash comes back identical and the replay costs zero tokens. That's the difference between a diff you can trust and a diff that's just noise.

Here's the whole thing: the architecture, the four design rules that make determinism provable, and the handful of traps I walked into so you don't have to.


The shape of it

browser (Vite + React 18 + TS + Tailwind)
  Graph · NodeEditor · DiffPanel · CostBar · RawJSON · Settings
        │  /api  (Vite proxy)
        ▼
FastAPI ──── spawns a daemon thread per run; client polls /graph for live nodes
        │
        ▼
Burr application ── plan → research → analyse → critique → revise → compose → verify → publish
  │        (llm)     (tool)    (tool)     (llm)      (llm)    (llm)     (tool)    (tool)
  ├── SQLitePersister("burr_state")        ← state after every node; what forking reads
  └── NodeTelemetryHook(PostRunStepHook)   ← inputs/outputs/latency/tokens/content-hash per node
        │
        ▼
rewind.db (SQLite, WAL): runs · nodes · edges · cache · tool_cache · burr_state
data/*.csv, data/*.json  →  artifacts/*.md
Enter fullscreen mode Exit fullscreen mode

Eight nodes, four real model calls, four real local computations over a real dataset file. The graph is deliberately boring — a linear chain — because the interesting axis is time, not topology.


Burr already had forking. I just had to find the right door.

I expected to hand-roll "resume from step N". Burr 0.42 has it natively:

builder.initialize_from(
    persister,
    resume_at_next_action=True,               # ← the important one
    default_state=initial,
    default_entrypoint=node_id,
    fork_from_app_id=src_app_id,
    fork_from_sequence_id=src_sequence_id,
)
Enter fullscreen mode Exit fullscreen mode

The trap: to re-run node N, you fork from N's predecessor row. A completed row's successor becomes the entrypoint. Pass N's own row and you silently start at N+1 — which is exactly what my first spike did, and it looked like the framework was broken. It wasn't; I'd asked for the state after the node I wanted to re-run.

parent    ── plan ─ research ── analyse ─ critique ── … ── publish
rows        seq 0     seq 1       seq 2      seq 3            seq 7
                            └── fork_from_sequence_id = 1, resume_at_next_action=True

fork      ┈┈┈ inherited ┈┈┈▶ analyse ▶ critique ▶ … ▶ publish
          copied verbatim,      └── re-executed (cache hits in cache mode)
          inherited: true
Enter fullscreen mode Exit fullscreen mode

Forking the first node has no predecessor row, so that one case starts a fresh application carrying the same override. And nested forks work because each node row keeps the src_app_id and sequence_id of the app that genuinely produced it — so forking inside a fork resolves to real executed state rather than to a copy of a copy.


Rule 1: state is semantic only

This is the rule that makes byte-identical replay possible, and it's the one I'd violate first if I weren't watching.

Latency, token counts, costs and run ids never enter Burr state. They travel in the action's result payload and get written to the nodes table by a PostRunStepHook:

class NodeTelemetryHook(PostRunStepHook):
    def post_run_step(self, *, app_id, partition_key, sequence_id,
                      state, action, result, exception, **kw):
        t = (result or {}).get("_telemetry") or {}
        db.upsert_node(self.db_path, run_id=app_id, node_id=action.name,
                       latency_ms=t["latency_ms"], total_tokens=t["total_tokens"],
                       cost_usd=t["cost_usd"], cache_hit=t["cache_hit"],
                       output_sha256=t["output_sha256"], …)
Enter fullscreen mode Exit fullscreen mode

Think about why. Node 3's prompt embeds state_json. If node 2's latency lived in state, node 3's prompt would contain a wall-clock number, its cache key would differ on every run, and "replay the unchanged branch" would be a permanently broken feature — while still looking completely correct in the UI. The bug wouldn't surface as a crash; it would surface as nothing ever hits cache and you'd spend a week debugging a determinism claim that was structurally impossible.

Rule 2: strip the framework's own keys

Burr keeps __PRIOR_STEP and __SEQUENCE_ID inside persisted state. A child run's sequence ids differ from its parent's by construction, so they must never reach a prompt or a hash:

def semantic_state(state: State) -> dict:
    return {k: v for k, v in state.data.items() if not k.startswith("__")}
Enter fullscreen mode Exit fullscreen mode

Rule 3: overrides live outside state

The fork's override is passed to the run context, not injected into state. If overrides were a state key, then a no-change replay would carry {"node_overrides": {}} where the parent had no such key — different state_json, different prompt, cold cache, and the headline feature quietly false. The override still lands in state where it belongs: it changes the node's output, and outputs are state.

Rule 4: hash the content, not the container

The artifact hash is the hash of the markdown bytes. The run-scoped filename is recorded on the run row and deliberately excluded from the hashed tool result:

hashed_result = {"tool": result["tool"], "bytes": result["bytes"], "written": True}
Enter fullscreen mode Exit fullscreen mode

I got this wrong first. The publish node stored its path in the tool result, so every run diverged at node 8 — and my diff dutifully reported first divergence: publish, which was true, meaningless, and cost me a verification cycle to notice.


The cache key

key = sha256( model | temperature | exact prompt | tool result )
Enter fullscreen mode Exit fullscreen mode

Length-prefixed, so the concatenation can't collide:

def _part(value: str) -> str:
    return f"{len(value.encode()):d}:{value}"

def llm_cache_key(*, model, temperature, prompt, tool_result):
    return hash_text(_part(model) + "|" + _part(repr(float(temperature)))
                     + "|" + _part(prompt) + "|" + _part(tool_result or ""))
Enter fullscreen mode Exit fullscreen mode

Tool nodes get their own table keyed on (tool name, canonical args). That's not decoration: without it, an unchanged replay would recompute every tool step and get the same answer by luck rather than by proof — and "every node is a cache hit" would be a claim about arithmetic, not about the system.

Each node also records why it was or wasn't a hit, because a boolean can't tell the truth here:

cache_source meaning
live went to the provider / recomputed; tokens charged
cache LLM response replayed from cache
tool-cache tool result replayed from the tool cache
override value supplied by the fork; nothing executed
render-cache artifact bytes unchanged (the file is still written for this run)
inherited copied from the parent; never executed

The diff engine: name the field, not the blob

Comparing two nested dicts and reporting "these differ" is useless. So flatten to leaves:

def flatten(value, prefix=""):
    if isinstance(value, dict):
        for key, item in value.items():
            yield from flatten(item, f"{prefix}.{key}" if prefix else str(key))
    elif isinstance(value, list):
        for index, item in enumerate(value):
            yield from flatten(item, f"{prefix}[{index}]")
    else:
        yield prefix or ".", value
Enter fullscreen mode Exit fullscreen mode

Then the diff says:

first divergence: analyse → tool result JSON (analyse tool result)
  result.groups.A.cost_per_outcome   26.4682  →  79.4046
  result.groups.A.total_cost         344.75   →  1034.25
artifact  0f661db5… → af9df990…   differs: true
Enter fullscreen mode Exit fullscreen mode

and the two published artifacts really do read differently — the parent says A | 26.47, the fork says Policy A: … 79.4046. That's the whole point of the tool: a one-field counterfactual, visible in prose.


Fail loudly, three times over

The brief's guardrail was "on parse failure the node fails loudly with the raw output attached; the run is marked failed, not silently patched". I kept that and added two more, because silent tolerance is how a determinism claim rots:

  1. Unparseable / missing / empty output → NodeFailure(node, msg, raw=raw). No fence-stripping retry loop that manufactures a passing answer.
  2. An override the node can't consume. A prompt edit on a tool node has no prompt to edit. Originally my code just… ignored it, and the fork produced an identical artifact. The test failed with first divergence: publish, and the real bug was that I'd applied a no-op. Now it raises, and the editor only offers fields the node kind actually uses.
  3. A client-supplied key is used for that request and never persisted, echoed, or returned — /api/settings exposes hasKey: true and nothing else. Not even a masked prefix; I had one and removed it, because "first four characters" is still key material on the wire.

What verification proved

npm run verify boots its own server on an OS-assigned free port and asserts, per task: 7+ nodes with hashed outputs and a real artifact file; fork at node 3 with inherited 1–2 and re-executed 3+; the diff naming node 3 and the exact field; an unchanged replay hitting 100% cache with 0 incremental tokens; and a different artifact hash for the fork. 14/14, exit 0.

The number I like most: the parent pass cost 10,054 tokens; the unchanged replay cost 0, and its artifact file is byte-identical to the fork's (cmp clean). I also recomputed the CSV maths independently, outside the app, and it matches the artifact exactly — so "real computation" isn't taking my word for it.

Two findings from poking at it:

  • A run survived a dead provider. Pointed at an unreachable base URL, it completed — every LLM node was cached and the rest were local tool calls. Forcing a cache miss against the same dead endpoint failed instantly with the transport error attached. That's the determinism property behaving like a feature.
  • LM Studio silently substitutes an unknown model id. A model override to no-such-model-xyz returned HTTP 200 with model: google/gemma-4-e4b. So model overrides are only meaningful against a provider that validates the id. Rewind records the model the provider answered with, so the substitution is visible in the node record instead of hidden.

Traps I'd like a week of my life back

  • The OpenAI SDK and Particle's metadata. The gateway returns metadata as an object whenever tools/structured output are involved; the SDK types it str | None, so every call dies with UnexpectedModelBehavior. A plain httpx POST parses the body itself and ignores unknown fields — immune by omission. The spec asked for plain fetch anyway.
  • response_format: {"type":"json_object"} is accepted by Particle and rejected by LM Studio with a 400 naming the field. One targeted retry without it keeps both presets usable while the specified behaviour stays primary.
  • Port clusters of look-alike processes. Several sibling builds run uvicorn agent.main:app with identical command lines on 3001/3011/…. A 2xx from a port probe proves nothing about which app answered. So /api/health carries "service": "rewind" and verification asserts on it.
  • Vite's IPv6 default lets a second dev server "succeed" on :5173 while the first owns 127.0.0.1 — then curl 127.0.0.1:5173 reports the wrong app. Bind 127.0.0.1 explicitly and make the proxy target an env var.
  • My own test harness deleted a live run's output. The scratch mechanics script did rmtree("artifacts") at startup, and I ran it while a verification was mid-flight. Four checks failed with a confusing 404, and the cause was me, not the code. Now the artifact directory is an env override (REWIND_ARTIFACT_DIR) and the harness writes into .scratch/. Lesson: a test that cleans a shared directory is a test that will sabotage a parallel run.

What's next

The highest-leverage addition is a what-if grid: fork one node across N overrides in parallel and render an N-way matrix, so "which of these five retrieval queries changes the conclusion?" is one screen instead of five forks. After that: non-linear graphs (which makes the diff align two trees), streaming node updates over SSE instead of 1.2 s polling, and a hash-stability mode that runs the same branch k times in live mode and reports the distribution of output hashes — turning "is my agent deterministic?" into a number rather than a binary.


Try it

git clone <your-fork> && cd rewind
uv venv --python 3.12 .venv && uv pip install -p .venv/bin/python -r requirements.txt
cp .env.example .env      # set base URL / key / model
npm install --prefix client
npm run dev               # → API + client, on ports it proves are free
npm run verify            # → 14 checks, exit 0
Enter fullscreen mode Exit fullscreen mode

The README documents the fork semantics, the cache key, the data model, the API surface, and every substitution I made against the original brief — including the fact that the recorded verification ran on a local model because no provider key was present in this project's environment.

Feedback and PRs welcome.

Code & more: https://www.dailybuild.xyz/project/274-rewind

Top comments (0)