DEV Community

Bogdan
Bogdan

Posted on

Auditable agents: turn the answer into a claim you can check

A model tells you "Q2 cloud spend was $45,000 — a 12.5% variance over budget."
Now two uncomfortable questions:

  1. Why? Which spreadsheet cell, which policy clause, which calculation?
  2. Again? If you run it tomorrow, do you get the same answer — and can you prove it's the same?

Most agent frameworks can't answer either cleanly. The run is a stream of
messages; the "reasoning" is prose in a log; the numbers came from wherever the
model felt like. When something goes wrong in production, your audit trail is a
transcript you read by eye.

This is the second post in a short series about building agents whose answers are
checkable. The first one was about side effects (an outbox
so a replay doesn't re-send). This one is about the other half: making the state
behind an answer inspectable and reproducible. It's the approach behind
reactifact, but the ideas — typed
artifacts, typed edges, a content hash over state — are portable.

TL;DR — Make the answer a typed artifact in a versioned context, not a
message in a bag. Link each derivation to its inputs as typed edges. Then "why
did it say that?" is a graph walk and "is this the same run?" is a hash — not a
transcript you read by eye.

Runnable offline in about a minute, no API key (the figures are computed in
Python, nothing calls a model):

pip install reactifact
python -m examples.fintech_audit.main   # prints the audit report, re-hashes to prove it's reproducible
Enter fullscreen mode Exit fullscreen mode

Code: github.com/bzdvdn/reactifact

Typical agent framework reactifact
The answer is a string in a message log a typed artifact in the context
State a message list versioned commits you can diff
"Why?" read the transcript by eye walk typed provenance edges
"Same run?" trust the prompt and the model compare a context_hash

Treat the answer as data, not text

The core move is boring and it changes everything: every meaningful thing an
agent produces is a typed artifact in one evolving context
, not a message in a
bag.

Context v1   Question
Context v2   + Document, Document
Context v3   + Evidence, Evidence
Context v4   + Claim
Context v5   + VerifiedClaim
Context v6   + Answer
Enter fullscreen mode Exit fullscreen mode

An artifact is a pydantic model — Evidence(text=..., source=...,
locator="budget.csv")
, Variance(pct=0.125). It has an id, a version, a
content hash, and who produced it. The context is versioned like a git history:
every step is a commit you can diff, roll back, check out.

Already that's more than a transcript. But the interesting part is the edges.

Provenance is structure, not a log line

When a produce derives something, it doesn't write "used budget.csv" into the
prose. It records a typed relation — a first-class edge in the artifact
graph:

source = self.effects.create(SourceRef(locator="transactions.csv"), id="ref:tx")
table = self.effects.create(Table(rows=...), id="doc:transactions.csv")
spend = self.effects.create(Spend(total=45000.0), id="spend:q2")
variance = self.effects.create(Variance(pct=0.125), id="variance:q2")
answer = self.effects.create(AuditAnswer(text="...$45,000... (+12.5%)"), id="answer:q2")

answer.link("supported_by", variance)          # claim ← its calculation
variance.link("calculated_from", spend)        # calculation ← its inputs
spend.link("materialized_from", table)         # figure ← the materialized table
table.link("materialized_from", source)        # table ← the source it was read from
Enter fullscreen mode Exit fullscreen mode

The relations are queryable (context.related(answer.id, "supported_by")), so
"why did it say that?" becomes a graph walk, not a grep. And because the graph is
built, you can render it — Mermaid in the CLI/dashboard, or a structured report:

from reactifact.audit import build_report, report_to_markdown

report = build_report(context, answer)   # walks the whole chain, breadth-first
print(report_to_markdown(report))
Enter fullscreen mode Exit fullscreen mode

build_report returns every artifact that contributed to the answer — each with
its content hash, version, and producing author — plus the source locators
the answer rests on. That's the "why": a machine-checkable provenance chain,
not a paragraph you have to believe.

# Audit report
- context version: 7
- context sha256: `f38c6a42…`
- Answer (sha256 c1d0…, by "finalize")
  - supported_by → Variance (sha256 9a51…, by "compute_variance")
    - calculated_from → Spend (sha256 4f2c…, by "compute_spend")
      - materialized_from → Table (sha256 77b1…, locator "budget.csv")
Enter fullscreen mode Exit fullscreen mode

Reproducibility: hash the state, not the vibes

Provenance answers "why". Reproducibility answers "is this the same run".

Because the context is canonical, you can fingerprint it. context_hash is a
sha256 over the run's state — each artifact's id, type, version and content hash,
plus every relation edge. Timestamps are deliberately excluded, so two runs
that reach the same state hash identically:

from reactifact.audit import context_hash

first = context_hash(await run_pipeline())
second = context_hash(await run_pipeline())
assert first == second         # reproducible — or it fails loudly
Enter fullscreen mode Exit fullscreen mode

That single string is an audit primitive. Save it next to the answer; later, a
reviewer re-runs the pipeline (or replays a saved session) and compares:

reactifact replay sessions.sqlite3 --session q2 --verify f38c6a42…
# exits non-zero on any mismatch
Enter fullscreen mode Exit fullscreen mode

No more "the numbers look about right." The state behind the answer either
hashes to the recorded fingerprint or it doesn't.

Determinism is a discipline, not a property you get for free

A hash only means something if the inputs are controlled. An agent run has three
usual sources of nondeterminism, and each has a handle:

  • The model. Record its calls once with a replaying provider, then re-run offline without the network:
  from reactifact.replay import ReplayLLM

  # pass 1 — record a real run
  resources = RuntimeResources(llm=ReplayLLM("calls.jsonl", mode="record", inner=real_llm))
  # pass 2 — reproduce it exactly; a divergent call raises ReplayMiss, never guesses
  resources = RuntimeResources(llm=ReplayLLM("calls.jsonl", mode="replay"))
Enter fullscreen mode Exit fullscreen mode
  • Auto-generated ids (uuid4) and wall-clock time. Pass a deterministic id
    factory and stop seeding artifact data from time.time() / uuid4(). A
    recorded model plus counter_ids() is often the whole fix.

  • Order and set iteration. Prefer stable, content-derived ids and explicit
    sorting in your produces.

To catch a leak, run the pipeline a few times under a recorded model and strict
ids and compare the fingerprints:

from reactifact.replay import verify_run

report = await verify_run(build, recording="calls.jsonl")   # runs it twice
assert report.ok, report.hashes    # a diff is real nondeterminism in your code
Enter fullscreen mode Exit fullscreen mode

That's the difference between "it usually returns the same thing" and "a second
run hashes to the same string."

Replay is reconstruction, not re-execution

Here's the subtle part, and where this connects to the outbox from the first post.

A naive "replay" re-runs the agent — which means it can hit the network again,
call tools again, and drift. reactifact's replay instead rebuilds the state from
the commit chain without running any agent
:

from reactifact.replay import replay_context, replay_summary

context = await replay_context(store, session_id, version=7)   # state at commit 7
print(replay_summary(context))    # counts by artifact type, relations, actions
Enter fullscreen mode Exit fullscreen mode

Because the commit chain is deterministic, you can reconstruct the exact context
at any point — walk the provenance, render the graph, answer "why" — and, since
no agent runs, nothing external fires. Pair it with the outbox and a recorded
side effect is read back as state instead of being re-sent.

The same versioning makes alternative states cheap: context.branch() to
explore two hypotheses, three-way merge() with explicit conflicts (no silent
last-write-wins), context.diff(v4, v9) to see exactly what changed between two
turns, context.checkout(v7) to move head back and undo a bad step. Time-travel over one artifact
graph, not a checkpoint of a message list.

Evaluate the state, not just the string

Once the run is structured state, evaluation stops being answer ==
expected_answer
. You can score the layers separately:

Evidence quality · Claim correctness · Provenance grounding ·
Calculation correctness · Confidence calibration · Answer quality · Source coverage
Enter fullscreen mode Exit fullscreen mode

reactifact.eval runs multi-level metrics over the final Context — including
provenance grounding, i.e. "is the answer actually linked to evidence that
supports it?" — which is a structural check, not a model's opinion. And because
the state is reproducible, a metric that passes today passes on replay.

The trace is part of the audit, and it's structured too

The report answers "why" from the final state. The same design makes the
runtime trace worth keeping: a run isn't a wall of log lines you read by
eye, it's a directed record of what actually happened. Each agent span carries
the artifact type that triggered the agent, and which Produce(s) ran for that
event — with how many effect operations each authored and how long it took. So
"the model said X" decomposes into "this event woke this agent, and this
produce did the work", not a black box.

That view is built in, not bolted on: the local SQLite dashboard
(create_trace_router) shows a Consume → Produce flow on each span, and the
same spans go to Langfuse (one child observation per produce, so the waterfall
shows each step) or any OTLP collector through one Tracer. Audit becomes two
views of one thing — the final provenance graph you can hash, and the causal
trace that built it.

The run trace in the local dashboard: work grouped by agent, each span a Consume → Produce flow with authored-operation counts, plus a live sequence diagram of every artifact write and LLM call.

The honest limits

Auditability here is a property of your pipeline, and it only holds as far as
you make it hold:

  • Determinism is earned. context_hash excludes timestamps, but if your produce calls uuid4() or reads the clock into artifact data, two runs will differ — verify_run tells you, it doesn't fix it.
  • It's single-process. Versioned state and replay are per context; this is not a distributed ledger.
  • Provenance is only as good as your links. The framework gives you typed edges and a report; if a produce doesn't link its evidence, the audit shows a gap — honestly, rather than inventing a justification.

That's the wager: an answer should be a claim you can check, and the machinery to
check it — typed artifacts, typed edges, a content hash over state, replay that
reconstructs — is worth building into the framework rather than bolting onto the
logs afterwards.

Run it offline

The fintech_audit example is exactly the scenario above — a variance over two
CSVs and a policy doc — with no API key (nothing calls a model; the figures are
computed in Python):

.venv/bin/python -m examples.fintech_audit.main
Enter fullscreen mode Exit fullscreen mode

It prints the computed figures, the audit report with a content hash per
artifact, and then runs the pipeline again and asserts the two hashes match — its
exit code is a determinism smoke test.

If you've shipped agents you had to debug at 2am: what's your audit trail —
logs you read by eye, or state you can query and reproduce?
I'd like to hear
where the first approach broke for you.

Top comments (0)