Your agent gave one answer yesterday and a different one today. Same code, same input file, a different decision. Before you can fix anything you need to know where the two runs parted ways. Did the prompt change, did the model reply change, or did something after the model call change?
This tutorial is for engineers who run agents in Python and want to answer that from two recorded runs instead of from memory. You will capture two runs as Run Capsules, compare them with nova diff, read which model call changed, and turn the same comparison into a CI gate. Every command and output below was captured from the released novafabric 0.104.0. Long paths and skipped lines are marked with ….
What you need
- Python 3.12 or newer (the package declares
Requires-Python: >=3.12) pip install novafabric openai- The
examples/blackbox_demo/directory from the public MSKazemi/novafabric repository, copied into an empty working directory. Run every command from that working directory.
No API key and no hosted model are involved. The example ships a small OpenAI-compatible mock server on 127.0.0.1:9099.
The scenario
agent.py reads a payment-service config, asks a model for one recommended change, and writes the answer to outputs/decision.json. The mock server answers in one of two ways, chosen by a request header the agent sends: a risky answer (disable rate limiting) or a safe one (reduce max_connections). That gives you two runs whose behaviour differs for a known reason, so you can check that the diff points at the right place.
Step 1: start the mock model
python blackbox_demo/mock_llm_server.py &
export OPENAI_API_KEY=sk-demo-no-key-needed
export OPENAI_BASE_URL=http://127.0.0.1:9099
export NOVAFABRIC_SUGGEST=0
Step 2: capture the first run
nova capture wraps the command with no code changes. Model calls from Python workloads are recorded automatically.
$ nova capture -- python blackbox_demo/agent.py --mode bad
✓ Capsule written: …/.novafabric/capsules/01M4FBTQ7FVQEZBWQGFD2RVTW2
Keep the capsule path in a variable (your run IDs will differ):
BAD=$HOME/.novafabric/capsules/01M4FBTQ7FVQEZBWQGFD2RVTW2
Step 3: capture the second run
$ nova capture -- python blackbox_demo/agent.py --mode fixed
✓ Capsule written: …/.novafabric/capsules/01M4FBVEZDG7XVBYJBB9WD6Q8D
Step 4: diff the two runs
$ nova diff "$BAD" "$FIXED"
changed=3 added=0 removed=0
Model calls:
~ call at span 5fef2505ba0bd78e
Outputs:
~ outputs/decision.json
~ outputs/stdout.txt
Three items changed (~): one model call and two output files. Nothing was added or removed. For the machine-readable record, add --output-format json; the model-call pair reports request_changed: false and response_changed: true.
How to read it
-
aligned: 1: the diff paired the one model call in each run. -
request_changed: false,response_changed: true: the agent sent the same request both times and got a different reply. The runs diverged at the model's answer, not in your prompt-building code. -
environment.changes: []: no recorded environment difference to chase. -
outputs: both files changed downstream of that reply.
Step 5: use it as a CI gate
--assert-no-regressions makes nova diff exit 1 when it finds any structural change between the two runs, including added or removed calls. Comparing a run with itself exits 0. On GitHub Actions, --output-format github-annotation prints one annotation per change.
Read the gate for what it is: it answers "did anything change?", not "did it get worse?". A safer answer fails it just as a riskier one does, as it did here.
What this does not do
- It does not judge the change. The diff is structural: it tells you which call and which files differ, not whether the new behaviour is better.
- Any difference counts. Against a live model, two runs of the same prompt can return different text, and the gate fails on that too.
-
Automatic model-call capture is for Python workloads. Other clients go through
nova api-proxy. - Sealing is opt-in and was not used here. The capsule format is pre-1.0.
The full version, with every command output, the replay section and the CI examples, is on the original page: https://novafabric.ai/blog/find-where-two-agent-runs-diverged/
Originally published at novafabric.ai. NovaFabric is open source (Apache-2.0) and pre-1.0: github.com/MSKazemi/novafabric.
Top comments (0)