<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Utsab Dahal</title>
    <description>The latest articles on DEV Community by Utsab Dahal (@utsab_dahal_0b150f9b4f42d).</description>
    <link>https://dev.to/utsab_dahal_0b150f9b4f42d</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F3967451%2F44532d0c-aeb8-47d9-afaf-5195ca540b78.jpg</url>
      <title>DEV Community: Utsab Dahal</title>
      <link>https://dev.to/utsab_dahal_0b150f9b4f42d</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/utsab_dahal_0b150f9b4f42d"/>
    <language>en</language>
    <item>
      <title>Your LLM Was Right. Your AI Agent Still Shipped the Wrong Answer.</title>
      <dc:creator>Utsab Dahal</dc:creator>
      <pubDate>Sat, 10 Oct 2026 05:58:45 +0000</pubDate>
      <link>https://dev.to/utsab_dahal_0b150f9b4f42d/your-llm-was-right-your-ai-agent-still-shipped-the-wrong-answer-25h4</link>
      <guid>https://dev.to/utsab_dahal_0b150f9b4f42d/your-llm-was-right-your-ai-agent-still-shipped-the-wrong-answer-25h4</guid>
      <description>&lt;p&gt;&lt;em&gt;I built an open-source Python tool to record AI agent runs, replay failures offline, and turn them into pytest regression tests. Here's what happened when I tested it with a real Gemini model.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;GitHub:&lt;/strong&gt; &lt;a href="https://github.com/utsab345/stepfork" rel="noopener noreferrer"&gt;https://github.com/utsab345/stepfork&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Real-LLM case study:&lt;/strong&gt; &lt;a href="https://github.com/utsab345/stepfork/tree/main/examples/real-llm-case-study" rel="noopener noreferrer"&gt;https://github.com/utsab345/stepfork/tree/main/examples/real-llm-case-study&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  The problem with debugging AI agents
&lt;/h2&gt;

&lt;p&gt;When a traditional Python function fails, reproducing the issue is usually straightforward. You provide the same input, run the function again, and investigate what went wrong.&lt;/p&gt;

&lt;p&gt;With AI agents, things get more complicated.&lt;/p&gt;

&lt;p&gt;An agent might call an LLM, retrieve documents, execute several tools, and make decisions based on intermediate results.&lt;/p&gt;

&lt;p&gt;When something goes wrong, running the agent again might produce a different result. The model can respond differently, external data can change, and tools might have side effects.&lt;/p&gt;

&lt;p&gt;I kept thinking about a simple question:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;What if we could record an agent failure once and turn it into a normal regression test?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;That's why I started building &lt;strong&gt;Stepfork&lt;/strong&gt;, an open-source Python library for recording and replaying AI agent executions.&lt;/p&gt;

&lt;p&gt;The idea is simple:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Record an agent execution, including instrumented LLM and tool interactions.&lt;/li&gt;
&lt;li&gt;Save those interactions in a portable trace.&lt;/li&gt;
&lt;li&gt;Replay the recorded interactions without calling the model again.&lt;/li&gt;
&lt;li&gt;Generate a pytest regression test.&lt;/li&gt;
&lt;li&gt;Fix the application and verify the behavior against the same trace.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;But I wanted to test this with a real model, not just mocked responses.&lt;/p&gt;

&lt;h2&gt;
  
  
  A real Gemini experiment
&lt;/h2&gt;

&lt;p&gt;I built a small IT incident-triage agent.&lt;/p&gt;

&lt;p&gt;It receives an incident report, checks service health and recent deployments, reads a runbook, and determines how serious the incident is.&lt;/p&gt;

&lt;p&gt;For the experiment, I used this synthetic incident:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;A production API is returning HTTP 500 errors for most customers. Error rates increased sharply after a deployment, and the payment checkout endpoint is failing.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;The agent used three types of local tools:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;code&gt;lookup_service_health&lt;/code&gt;&lt;/li&gt;
&lt;li&gt;&lt;code&gt;get_recent_deployments&lt;/code&gt;&lt;/li&gt;
&lt;li&gt;&lt;code&gt;fetch_incident_runbook&lt;/code&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The operational data was synthetic, but the LLM request was real.&lt;/p&gt;

&lt;p&gt;I used &lt;strong&gt;Gemini 2.5 Flash&lt;/strong&gt;, accessed through Google's OpenAI-compatible API endpoint.&lt;/p&gt;

&lt;p&gt;The local fixtures indicated that the checkout API had a 72% error rate, customers were affected, and a recent deployment had occurred.&lt;/p&gt;

&lt;p&gt;According to the independently defined incident-severity policy, this was a &lt;strong&gt;P0 incident requiring immediate escalation&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;And Gemini got it right.&lt;/p&gt;

&lt;p&gt;The model returned:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"severity"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"P0"&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Then my Python code changed it to P1.&lt;/p&gt;

&lt;h2&gt;
  
  
  The bug wasn't in the model
&lt;/h2&gt;

&lt;p&gt;The agent had a deterministic postprocessing function that correlated incidents with recent deployments.&lt;/p&gt;

&lt;p&gt;I intentionally introduced a bug into this function:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;_correlate_recent_deployment&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;decision&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;evidence&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="n"&gt;recent&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;evidence&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;deployments&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;][&lt;/span&gt;&lt;span class="n"&gt;INCIDENT&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;primary_service&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;]][&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;recent_deployment&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;

    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;recent&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="n"&gt;decision&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;severity&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;P1&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
        &lt;span class="n"&gt;decision&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;escalation_required&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="bp"&gt;False&lt;/span&gt;

    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;decision&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The code incorrectly assumed that a recent deployment justified lowering the incident severity.&lt;/p&gt;

&lt;p&gt;So the final output became:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"severity"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"P1"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"escalation_required"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="kc"&gt;false&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The model had correctly identified a critical incident, but the application silently overrode its decision.&lt;/p&gt;

&lt;p&gt;No exception was raised. The agent completed successfully.&lt;/p&gt;

&lt;p&gt;A test that only checked whether the agent ran without crashing would have passed.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The failure was in the business outcome, not the execution status.&lt;/strong&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Recording the execution with Stepfork
&lt;/h2&gt;

&lt;p&gt;I used Stepfork to record the agent's instrumented dependencies.&lt;/p&gt;

&lt;p&gt;The resulting &lt;code&gt;.sftrace&lt;/code&gt; bundle contained:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Recorded information&lt;/th&gt;
&lt;th&gt;Result&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Model&lt;/td&gt;
&lt;td&gt;Gemini 2.5 Flash&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Total events&lt;/td&gt;
&lt;td&gt;14&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Tool calls&lt;/td&gt;
&lt;td&gt;5&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;LLM calls&lt;/td&gt;
&lt;td&gt;1&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Trace validation&lt;/td&gt;
&lt;td&gt;Passed&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Integrity verification&lt;/td&gt;
&lt;td&gt;Verified&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The trace captured the tool-call sequence, recorded inputs and outputs, and the LLM interaction.&lt;/p&gt;

&lt;p&gt;The recorded model response remained P0, while the application returned P1.&lt;/p&gt;

&lt;p&gt;This distinction is important because it tells us exactly where the behavior went wrong.&lt;/p&gt;

&lt;h2&gt;
  
  
  Replaying the failure without calling Gemini again
&lt;/h2&gt;

&lt;p&gt;Next, I used Stepfork's frozen replay mode.&lt;/p&gt;

&lt;p&gt;Instead of executing the instrumented tools or sending another request to Gemini, Stepfork substituted the recorded responses.&lt;/p&gt;

&lt;p&gt;The application logic still executed, so I could change that logic and test the outcome.&lt;/p&gt;

&lt;p&gt;The verification results were:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Instrumented tool bodies executed: 0
Live LLM attempts:              0
Recorded dependency calls:      6
Substituted dependency calls:    6
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That means I could reproduce the relevant execution without additional model API calls.&lt;/p&gt;

&lt;p&gt;It also meant the regression test could run offline, without an API key.&lt;/p&gt;

&lt;p&gt;There is an important limitation: Stepfork only substitutes dependencies that have been instrumented. It does not prevent arbitrary uninstrumented code from executing, and replay is not a sandbox.&lt;/p&gt;

&lt;h2&gt;
  
  
  Turning the failure into a pytest regression test
&lt;/h2&gt;

&lt;p&gt;I defined the expected business outcome independently:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"severity"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"P0"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"escalation_required"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="kc"&gt;true&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Then I exported a regression test using Stepfork:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;stepfork &lt;span class="nb"&gt;export &lt;/span&gt;traces/incident-triage.sftrace &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--pytest&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--entrypoint&lt;/span&gt; agent:run_incident_agent &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--expect-output&lt;/span&gt; expected.json &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--output&lt;/span&gt; tests/test_incident_regression.py
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Running the generated test against the buggy application produced:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;1 failed

AssertionError: behavior mismatch at 'escalation_required':
expected true, got false
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That was exactly what I wanted.&lt;/p&gt;

&lt;p&gt;The test detected a real application-level mistake, using the recorded model interaction.&lt;/p&gt;

&lt;h2&gt;
  
  
  Fixing the bug
&lt;/h2&gt;

&lt;p&gt;The fix was simple: stop letting deployment correlation override the incident-severity policy.&lt;/p&gt;

&lt;p&gt;The corrected function became:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;_correlate_recent_deployment&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;decision&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;evidence&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;decision&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;I ran the same generated pytest test again.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;1 passed
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;I didn't regenerate the trace.&lt;/p&gt;

&lt;p&gt;I didn't change the expected outcome.&lt;/p&gt;

&lt;p&gt;I didn't make another Gemini request.&lt;/p&gt;

&lt;p&gt;The only meaningful change was the application logic.&lt;/p&gt;

&lt;p&gt;This was the result:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;&lt;/th&gt;
&lt;th&gt;Before fix&lt;/th&gt;
&lt;th&gt;After fix&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Recorded model decision&lt;/td&gt;
&lt;td&gt;P0&lt;/td&gt;
&lt;td&gt;P0&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Application decision&lt;/td&gt;
&lt;td&gt;P1&lt;/td&gt;
&lt;td&gt;P0&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Escalation&lt;/td&gt;
&lt;td&gt;False&lt;/td&gt;
&lt;td&gt;True&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Regression test&lt;/td&gt;
&lt;td&gt;Failed&lt;/td&gt;
&lt;td&gt;Passed&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Additional LLM requests&lt;/td&gt;
&lt;td&gt;0&lt;/td&gt;
&lt;td&gt;0&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;One nuance: replaying the fixed application with the standalone CLI can report output divergence from the originally recorded buggy result. That is expected. The generated pytest test separately verifies the corrected result against the independently defined expected outcome.&lt;/p&gt;

&lt;h2&gt;
  
  
  What happens when the agent's tool calls change?
&lt;/h2&gt;

&lt;p&gt;I also tested changes to the agent's dependency trajectory.&lt;/p&gt;

&lt;p&gt;A harmless refactor that preserved the tool calls, arguments, order, and model prompt passed.&lt;/p&gt;

&lt;p&gt;But when I changed the order of tool calls, frozen replay rejected the execution:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;ReplayMismatchError:
call #1: expected tool 'lookup_service_health'
but the agent called 'get_recent_deployments'
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Changing a tool argument from &lt;code&gt;window_minutes=60&lt;/code&gt; to &lt;code&gt;window_minutes=120&lt;/code&gt; also triggered a mismatch.&lt;/p&gt;

&lt;p&gt;This is useful because an agent regression isn't always about the final answer. Sometimes the agent starts calling different tools or sending different arguments.&lt;/p&gt;

&lt;h2&gt;
  
  
  How much did the experiment cost?
&lt;/h2&gt;

&lt;p&gt;I made three real Gemini requests:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;One request to record the original execution.&lt;/li&gt;
&lt;li&gt;One hybrid replay with the baseline prompt.&lt;/li&gt;
&lt;li&gt;One hybrid replay with a revised prompt.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The estimated API cost was approximately &lt;strong&gt;$0.009&lt;/strong&gt;, based on the reported token usage and pricing assumptions.&lt;/p&gt;

&lt;p&gt;Both hybrid evaluations returned P0. With only one sample per prompt, that doesn't establish whether either prompt is better.&lt;/p&gt;

&lt;p&gt;The important result was that the recorded execution could be replayed repeatedly without additional model requests.&lt;/p&gt;

&lt;h2&gt;
  
  
  Try it yourself
&lt;/h2&gt;

&lt;p&gt;The complete case study is included in the main Stepfork repository:&lt;/p&gt;

&lt;p&gt;&lt;a href="https://github.com/utsab345/stepfork/tree/main/examples/real-llm-case-study" rel="noopener noreferrer"&gt;https://github.com/utsab345/stepfork/tree/main/examples/real-llm-case-study&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;You can reproduce the offline tests without a Gemini API key.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;git clone https://github.com/utsab345/stepfork.git

&lt;span class="nb"&gt;cd &lt;/span&gt;stepfork/examples/real-llm-case-study

uv &lt;span class="nb"&gt;sync&lt;/span&gt; &lt;span class="nt"&gt;--locked&lt;/span&gt; &lt;span class="nt"&gt;--extra&lt;/span&gt; &lt;span class="nb"&gt;test

&lt;/span&gt;uv run python &lt;span class="nt"&gt;-m&lt;/span&gt; pytest &lt;span class="nt"&gt;-q&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;To verify the recorded trace:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;uv run stepfork validate &lt;span class="se"&gt;\&lt;/span&gt;
  traces/incident-triage.sftrace &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--verify-integrity&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;To verify frozen replay:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;uv run python scripts/verify_frozen.py
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The case study includes the original trace, synthetic fixtures, buggy and fixed application snapshots, generated pytest regression test, and recorded verification evidence.&lt;/p&gt;

&lt;h2&gt;
  
  
  What Stepfork doesn't do yet
&lt;/h2&gt;

&lt;p&gt;Stepfork is still an early alpha project.&lt;/p&gt;

&lt;p&gt;It doesn't automatically know whether an agent's answer is correct. You still need to define meaningful expectations.&lt;/p&gt;

&lt;p&gt;It also doesn't intercept every possible side effect. Frozen replay applies to supported, instrumented dependencies.&lt;/p&gt;

&lt;p&gt;The current OpenAI adapter has limitations, including unsupported streaming and some newer API surfaces. The project also supports a LangGraph integration, with additional integrations planned.&lt;/p&gt;

&lt;p&gt;I'm working toward a simple developer workflow:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;An agent fails once. You record it. You fix the code. That failure becomes a test you can run again.&lt;/strong&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  I'd love feedback
&lt;/h2&gt;

&lt;p&gt;I'm building Stepfork as an open-source project, and I'd especially like to hear from developers working with LLM agents, LangGraph, and Python testing tools.&lt;/p&gt;

&lt;p&gt;What types of agent failures are hardest for you to reproduce?&lt;/p&gt;

&lt;p&gt;Would recording tool calls and replaying them offline help your debugging workflow?&lt;/p&gt;

&lt;p&gt;If you'd like to try it or contribute:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;GitHub:&lt;/strong&gt; &lt;a href="https://github.com/utsab345/stepfork" rel="noopener noreferrer"&gt;https://github.com/utsab345/stepfork&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Documentation:&lt;/strong&gt; &lt;a href="https://utsab345.github.io/stepfork/" rel="noopener noreferrer"&gt;https://utsab345.github.io/stepfork/&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Real-Gemini case study:&lt;/strong&gt; &lt;a href="https://github.com/utsab345/stepfork/tree/main/examples/real-llm-case-study" rel="noopener noreferrer"&gt;https://github.com/utsab345/stepfork/tree/main/examples/real-llm-case-study&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;The agent failed. Make the failure a test.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>opensource</category>
      <category>testing</category>
      <category>python</category>
    </item>
  </channel>
</rss>
