Rewind: counterfactual replay for agent pipelines
Every agent trace tool I've used has the same ceiling. It shows me a waterfall of steps, I spot that step six produced something stupid, and then I'm stuck. To find out why — or whether it was the model, the retrieval, or the prompt — I have to re-run the whole pipeline and hope it lands in the same place. With a non-deterministic model it doesn't, so I can never separate my change from model jitter.
So I built the other thing: fork a completed run at any step, change exactly one input, replay only the downstream steps, and diff the two trajectories. Rewind does that, and — this is the part that took real care — when you replay a branch you didn't change, every output hash comes back identical and the replay costs zero tokens. That's the difference between a diff you can trust and a diff that's just noise.
Here's the whole thing: the architecture, the four design rules that make determinism provable, and the handful of traps I walked into so you don't have to.
The shape of it
browser (Vite + React 18 + TS + Tailwind)
Graph · NodeEditor · DiffPanel · CostBar · RawJSON · Settings
│ /api (Vite proxy)
▼
FastAPI ──── spawns a daemon thread per run; client polls /graph for live nodes
│
▼
Burr application ── plan → research → analyse → critique → revise → compose → verify → publish
│ (llm) (tool) (tool) (llm) (llm) (llm) (tool) (tool)
├── SQLitePersister("burr_state") ← state after every node; what forking reads
└── NodeTelemetryHook(PostRunStepHook) ← inputs/outputs/latency/tokens/content-hash per node
│
▼
rewind.db (SQLite, WAL): runs · nodes · edges · cache · tool_cache · burr_state
data/*.csv, data/*.json → artifacts/*.md
Eight nodes, four real model calls, four real local computations over a real dataset file. The graph is deliberately boring — a linear chain — because the interesting axis is time, not topology.
Burr already had forking. I just had to find the right door.
I expected to hand-roll "resume from step N". Burr 0.42 has it natively:
builder.initialize_from(
persister,
resume_at_next_action=True, # ← the important one
default_state=initial,
default_entrypoint=node_id,
fork_from_app_id=src_app_id,
fork_from_sequence_id=src_sequence_id,
)
The trap: to re-run node N, you fork from N's predecessor row. A completed row's successor becomes the entrypoint. Pass N's own row and you silently start at N+1 — which is exactly what my first spike did, and it looked like the framework was broken. It wasn't; I'd asked for the state after the node I wanted to re-run.
parent ── plan ─ research ── analyse ─ critique ── … ── publish
rows seq 0 seq 1 seq 2 seq 3 seq 7
└── fork_from_sequence_id = 1, resume_at_next_action=True
fork ┈┈┈ inherited ┈┈┈▶ analyse ▶ critique ▶ … ▶ publish
copied verbatim, └── re-executed (cache hits in cache mode)
inherited: true
Forking the first node has no predecessor row, so that one case starts a fresh application carrying the same override. And nested forks work because each node row keeps the src_app_id and sequence_id of the app that genuinely produced it — so forking inside a fork resolves to real executed state rather than to a copy of a copy.
Rule 1: state is semantic only
This is the rule that makes byte-identical replay possible, and it's the one I'd violate first if I weren't watching.
Latency, token counts, costs and run ids never enter Burr state. They travel in the action's result payload and get written to the nodes table by a PostRunStepHook:
class NodeTelemetryHook(PostRunStepHook):
def post_run_step(self, *, app_id, partition_key, sequence_id,
state, action, result, exception, **kw):
t = (result or {}).get("_telemetry") or {}
db.upsert_node(self.db_path, run_id=app_id, node_id=action.name,
latency_ms=t["latency_ms"], total_tokens=t["total_tokens"],
cost_usd=t["cost_usd"], cache_hit=t["cache_hit"],
output_sha256=t["output_sha256"], …)
Think about why. Node 3's prompt embeds state_json. If node 2's latency lived in state, node 3's prompt would contain a wall-clock number, its cache key would differ on every run, and "replay the unchanged branch" would be a permanently broken feature — while still looking completely correct in the UI. The bug wouldn't surface as a crash; it would surface as nothing ever hits cache and you'd spend a week debugging a determinism claim that was structurally impossible.
Rule 2: strip the framework's own keys
Burr keeps __PRIOR_STEP and __SEQUENCE_ID inside persisted state. A child run's sequence ids differ from its parent's by construction, so they must never reach a prompt or a hash:
def semantic_state(state: State) -> dict:
return {k: v for k, v in state.data.items() if not k.startswith("__")}
Rule 3: overrides live outside state
The fork's override is passed to the run context, not injected into state. If overrides were a state key, then a no-change replay would carry {"node_overrides": {}} where the parent had no such key — different state_json, different prompt, cold cache, and the headline feature quietly false. The override still lands in state where it belongs: it changes the node's output, and outputs are state.
Rule 4: hash the content, not the container
The artifact hash is the hash of the markdown bytes. The run-scoped filename is recorded on the run row and deliberately excluded from the hashed tool result:
hashed_result = {"tool": result["tool"], "bytes": result["bytes"], "written": True}
I got this wrong first. The publish node stored its path in the tool result, so every run diverged at node 8 — and my diff dutifully reported first divergence: publish, which was true, meaningless, and cost me a verification cycle to notice.
The cache key
key = sha256( model | temperature | exact prompt | tool result )
Length-prefixed, so the concatenation can't collide:
def _part(value: str) -> str:
return f"{len(value.encode()):d}:{value}"
def llm_cache_key(*, model, temperature, prompt, tool_result):
return hash_text(_part(model) + "|" + _part(repr(float(temperature)))
+ "|" + _part(prompt) + "|" + _part(tool_result or ""))
Tool nodes get their own table keyed on (tool name, canonical args). That's not decoration: without it, an unchanged replay would recompute every tool step and get the same answer by luck rather than by proof — and "every node is a cache hit" would be a claim about arithmetic, not about the system.
Each node also records why it was or wasn't a hit, because a boolean can't tell the truth here:
cache_source |
meaning |
|---|---|
live |
went to the provider / recomputed; tokens charged |
cache |
LLM response replayed from cache |
tool-cache |
tool result replayed from the tool cache |
override |
value supplied by the fork; nothing executed |
render-cache |
artifact bytes unchanged (the file is still written for this run) |
inherited |
copied from the parent; never executed |
The diff engine: name the field, not the blob
Comparing two nested dicts and reporting "these differ" is useless. So flatten to leaves:
def flatten(value, prefix=""):
if isinstance(value, dict):
for key, item in value.items():
yield from flatten(item, f"{prefix}.{key}" if prefix else str(key))
elif isinstance(value, list):
for index, item in enumerate(value):
yield from flatten(item, f"{prefix}[{index}]")
else:
yield prefix or ".", value
Then the diff says:
first divergence: analyse → tool result JSON (analyse tool result)
result.groups.A.cost_per_outcome 26.4682 → 79.4046
result.groups.A.total_cost 344.75 → 1034.25
artifact 0f661db5… → af9df990… differs: true
and the two published artifacts really do read differently — the parent says A | 26.47, the fork says Policy A: … 79.4046. That's the whole point of the tool: a one-field counterfactual, visible in prose.
Fail loudly, three times over
The brief's guardrail was "on parse failure the node fails loudly with the raw output attached; the run is marked failed, not silently patched". I kept that and added two more, because silent tolerance is how a determinism claim rots:
-
Unparseable / missing / empty
output→NodeFailure(node, msg, raw=raw). No fence-stripping retry loop that manufactures a passing answer. -
An override the node can't consume. A
promptedit on a tool node has no prompt to edit. Originally my code just… ignored it, and the fork produced an identical artifact. The test failed withfirst divergence: publish, and the real bug was that I'd applied a no-op. Now it raises, and the editor only offers fields the node kind actually uses. -
A client-supplied key is used for that request and never persisted, echoed, or returned —
/api/settingsexposeshasKey: trueand nothing else. Not even a masked prefix; I had one and removed it, because "first four characters" is still key material on the wire.
What verification proved
npm run verify boots its own server on an OS-assigned free port and asserts, per task: 7+ nodes with hashed outputs and a real artifact file; fork at node 3 with inherited 1–2 and re-executed 3+; the diff naming node 3 and the exact field; an unchanged replay hitting 100% cache with 0 incremental tokens; and a different artifact hash for the fork. 14/14, exit 0.
The number I like most: the parent pass cost 10,054 tokens; the unchanged replay cost 0, and its artifact file is byte-identical to the fork's (cmp clean). I also recomputed the CSV maths independently, outside the app, and it matches the artifact exactly — so "real computation" isn't taking my word for it.
Two findings from poking at it:
- A run survived a dead provider. Pointed at an unreachable base URL, it completed — every LLM node was cached and the rest were local tool calls. Forcing a cache miss against the same dead endpoint failed instantly with the transport error attached. That's the determinism property behaving like a feature.
-
LM Studio silently substitutes an unknown model id. A
modeloverride tono-such-model-xyzreturned HTTP 200 withmodel: google/gemma-4-e4b. So model overrides are only meaningful against a provider that validates the id. Rewind records the model the provider answered with, so the substitution is visible in the node record instead of hidden.
Traps I'd like a week of my life back
-
The OpenAI SDK and Particle's
metadata. The gateway returnsmetadataas an object whenever tools/structured output are involved; the SDK types itstr | None, so every call dies withUnexpectedModelBehavior. A plainhttpxPOST parses the body itself and ignores unknown fields — immune by omission. The spec asked for plain fetch anyway. -
response_format: {"type":"json_object"}is accepted by Particle and rejected by LM Studio with a 400 naming the field. One targeted retry without it keeps both presets usable while the specified behaviour stays primary. -
Port clusters of look-alike processes. Several sibling builds run
uvicorn agent.main:appwith identical command lines on 3001/3011/…. A 2xx from a port probe proves nothing about which app answered. So/api/healthcarries"service": "rewind"and verification asserts on it. -
Vite's IPv6 default lets a second dev server "succeed" on :5173 while the first owns
127.0.0.1— thencurl 127.0.0.1:5173reports the wrong app. Bind127.0.0.1explicitly and make the proxy target an env var. -
My own test harness deleted a live run's output. The scratch mechanics script did
rmtree("artifacts")at startup, and I ran it while a verification was mid-flight. Four checks failed with a confusing 404, and the cause was me, not the code. Now the artifact directory is an env override (REWIND_ARTIFACT_DIR) and the harness writes into.scratch/. Lesson: a test that cleans a shared directory is a test that will sabotage a parallel run.
What's next
The highest-leverage addition is a what-if grid: fork one node across N overrides in parallel and render an N-way matrix, so "which of these five retrieval queries changes the conclusion?" is one screen instead of five forks. After that: non-linear graphs (which makes the diff align two trees), streaming node updates over SSE instead of 1.2 s polling, and a hash-stability mode that runs the same branch k times in live mode and reports the distribution of output hashes — turning "is my agent deterministic?" into a number rather than a binary.
Try it
git clone <your-fork> && cd rewind
uv venv --python 3.12 .venv && uv pip install -p .venv/bin/python -r requirements.txt
cp .env.example .env # set base URL / key / model
npm install --prefix client
npm run dev # → API + client, on ports it proves are free
npm run verify # → 14 checks, exit 0
The README documents the fork semantics, the cache key, the data model, the API surface, and every substitution I made against the original brief — including the fact that the recorded verification ran on a local model because no provider key was present in this project's environment.
Feedback and PRs welcome.
Code & more: https://www.dailybuild.xyz/project/274-rewind
Top comments (0)