I built an open-source Python tool to record AI agent runs, replay failures offline, and turn them into pytest regression tests. Here's what happened when I tested it with a real Gemini model.
GitHub: https://github.com/utsab345/stepfork
Real-LLM case study: https://github.com/utsab345/stepfork/tree/main/examples/real-llm-case-study
The problem with debugging AI agents
When a traditional Python function fails, reproducing the issue is usually straightforward. You provide the same input, run the function again, and investigate what went wrong.
With AI agents, things get more complicated.
An agent might call an LLM, retrieve documents, execute several tools, and make decisions based on intermediate results.
When something goes wrong, running the agent again might produce a different result. The model can respond differently, external data can change, and tools might have side effects.
I kept thinking about a simple question:
What if we could record an agent failure once and turn it into a normal regression test?
That's why I started building Stepfork, an open-source Python library for recording and replaying AI agent executions.
The idea is simple:
- Record an agent execution, including instrumented LLM and tool interactions.
- Save those interactions in a portable trace.
- Replay the recorded interactions without calling the model again.
- Generate a pytest regression test.
- Fix the application and verify the behavior against the same trace.
But I wanted to test this with a real model, not just mocked responses.
A real Gemini experiment
I built a small IT incident-triage agent.
It receives an incident report, checks service health and recent deployments, reads a runbook, and determines how serious the incident is.
For the experiment, I used this synthetic incident:
A production API is returning HTTP 500 errors for most customers. Error rates increased sharply after a deployment, and the payment checkout endpoint is failing.
The agent used three types of local tools:
lookup_service_healthget_recent_deploymentsfetch_incident_runbook
The operational data was synthetic, but the LLM request was real.
I used Gemini 2.5 Flash, accessed through Google's OpenAI-compatible API endpoint.
The local fixtures indicated that the checkout API had a 72% error rate, customers were affected, and a recent deployment had occurred.
According to the independently defined incident-severity policy, this was a P0 incident requiring immediate escalation.
And Gemini got it right.
The model returned:
{
"severity": "P0"
}
Then my Python code changed it to P1.
The bug wasn't in the model
The agent had a deterministic postprocessing function that correlated incidents with recent deployments.
I intentionally introduced a bug into this function:
def _correlate_recent_deployment(decision, evidence):
recent = evidence["deployments"][INCIDENT["primary_service"]]["recent_deployment"]
if recent:
decision["severity"] = "P1"
decision["escalation_required"] = False
return decision
The code incorrectly assumed that a recent deployment justified lowering the incident severity.
So the final output became:
{
"severity": "P1",
"escalation_required": false
}
The model had correctly identified a critical incident, but the application silently overrode its decision.
No exception was raised. The agent completed successfully.
A test that only checked whether the agent ran without crashing would have passed.
The failure was in the business outcome, not the execution status.
Recording the execution with Stepfork
I used Stepfork to record the agent's instrumented dependencies.
The resulting .sftrace bundle contained:
| Recorded information | Result |
|---|---|
| Model | Gemini 2.5 Flash |
| Total events | 14 |
| Tool calls | 5 |
| LLM calls | 1 |
| Trace validation | Passed |
| Integrity verification | Verified |
The trace captured the tool-call sequence, recorded inputs and outputs, and the LLM interaction.
The recorded model response remained P0, while the application returned P1.
This distinction is important because it tells us exactly where the behavior went wrong.
Replaying the failure without calling Gemini again
Next, I used Stepfork's frozen replay mode.
Instead of executing the instrumented tools or sending another request to Gemini, Stepfork substituted the recorded responses.
The application logic still executed, so I could change that logic and test the outcome.
The verification results were:
Instrumented tool bodies executed: 0
Live LLM attempts: 0
Recorded dependency calls: 6
Substituted dependency calls: 6
That means I could reproduce the relevant execution without additional model API calls.
It also meant the regression test could run offline, without an API key.
There is an important limitation: Stepfork only substitutes dependencies that have been instrumented. It does not prevent arbitrary uninstrumented code from executing, and replay is not a sandbox.
Turning the failure into a pytest regression test
I defined the expected business outcome independently:
{
"severity": "P0",
"escalation_required": true
}
Then I exported a regression test using Stepfork:
stepfork export traces/incident-triage.sftrace \
--pytest \
--entrypoint agent:run_incident_agent \
--expect-output expected.json \
--output tests/test_incident_regression.py
Running the generated test against the buggy application produced:
1 failed
AssertionError: behavior mismatch at 'escalation_required':
expected true, got false
That was exactly what I wanted.
The test detected a real application-level mistake, using the recorded model interaction.
Fixing the bug
The fix was simple: stop letting deployment correlation override the incident-severity policy.
The corrected function became:
def _correlate_recent_deployment(decision, evidence):
return decision
I ran the same generated pytest test again.
1 passed
I didn't regenerate the trace.
I didn't change the expected outcome.
I didn't make another Gemini request.
The only meaningful change was the application logic.
This was the result:
| Before fix | After fix | |
|---|---|---|
| Recorded model decision | P0 | P0 |
| Application decision | P1 | P0 |
| Escalation | False | True |
| Regression test | Failed | Passed |
| Additional LLM requests | 0 | 0 |
One nuance: replaying the fixed application with the standalone CLI can report output divergence from the originally recorded buggy result. That is expected. The generated pytest test separately verifies the corrected result against the independently defined expected outcome.
What happens when the agent's tool calls change?
I also tested changes to the agent's dependency trajectory.
A harmless refactor that preserved the tool calls, arguments, order, and model prompt passed.
But when I changed the order of tool calls, frozen replay rejected the execution:
ReplayMismatchError:
call #1: expected tool 'lookup_service_health'
but the agent called 'get_recent_deployments'
Changing a tool argument from window_minutes=60 to window_minutes=120 also triggered a mismatch.
This is useful because an agent regression isn't always about the final answer. Sometimes the agent starts calling different tools or sending different arguments.
How much did the experiment cost?
I made three real Gemini requests:
- One request to record the original execution.
- One hybrid replay with the baseline prompt.
- One hybrid replay with a revised prompt.
The estimated API cost was approximately $0.009, based on the reported token usage and pricing assumptions.
Both hybrid evaluations returned P0. With only one sample per prompt, that doesn't establish whether either prompt is better.
The important result was that the recorded execution could be replayed repeatedly without additional model requests.
Try it yourself
The complete case study is included in the main Stepfork repository:
https://github.com/utsab345/stepfork/tree/main/examples/real-llm-case-study
You can reproduce the offline tests without a Gemini API key.
git clone https://github.com/utsab345/stepfork.git
cd stepfork/examples/real-llm-case-study
uv sync --locked --extra test
uv run python -m pytest -q
To verify the recorded trace:
uv run stepfork validate \
traces/incident-triage.sftrace \
--verify-integrity
To verify frozen replay:
uv run python scripts/verify_frozen.py
The case study includes the original trace, synthetic fixtures, buggy and fixed application snapshots, generated pytest regression test, and recorded verification evidence.
What Stepfork doesn't do yet
Stepfork is still an early alpha project.
It doesn't automatically know whether an agent's answer is correct. You still need to define meaningful expectations.
It also doesn't intercept every possible side effect. Frozen replay applies to supported, instrumented dependencies.
The current OpenAI adapter has limitations, including unsupported streaming and some newer API surfaces. The project also supports a LangGraph integration, with additional integrations planned.
I'm working toward a simple developer workflow:
An agent fails once. You record it. You fix the code. That failure becomes a test you can run again.
I'd love feedback
I'm building Stepfork as an open-source project, and I'd especially like to hear from developers working with LLM agents, LangGraph, and Python testing tools.
What types of agent failures are hardest for you to reproduce?
Would recording tool calls and replaying them offline help your debugging workflow?
If you'd like to try it or contribute:
GitHub: https://github.com/utsab345/stepfork
Documentation: https://utsab345.github.io/stepfork/
Real-Gemini case study: https://github.com/utsab345/stepfork/tree/main/examples/real-llm-case-study
The agent failed. Make the failure a test.
Top comments (0)