<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: UJWAL BAGALKOTI</title>
    <description>The latest articles on DEV Community by UJWAL BAGALKOTI (@ujwal_bagalkoti).</description>
    <link>https://dev.to/ujwal_bagalkoti</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4156151%2F5cf876bd-9212-4fa9-b895-5c9dd00ac956.jpg</url>
      <title>DEV Community: UJWAL BAGALKOTI</title>
      <link>https://dev.to/ujwal_bagalkoti</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/ujwal_bagalkoti"/>
    <language>en</language>
    <item>
      <title>How I Built a Time-Travel Debugger for AI Agents</title>
      <dc:creator>UJWAL BAGALKOTI</dc:creator>
      <pubDate>Fri, 02 Oct 2026 03:50:29 +0000</pubDate>
      <link>https://dev.to/ujwal_bagalkoti/how-i-built-a-time-travel-debugger-for-ai-agents-b0h</link>
      <guid>https://dev.to/ujwal_bagalkoti/how-i-built-a-time-travel-debugger-for-ai-agents-b0h</guid>
      <description>&lt;h1&gt;
  
  
  How I Built a Time-Travel Debugger for AI Agents
&lt;/h1&gt;

&lt;p&gt;What if debugging an AI agent worked more like debugging normal code?&lt;/p&gt;

&lt;p&gt;Pause execution.&lt;/p&gt;

&lt;p&gt;Inspect what happened.&lt;/p&gt;

&lt;p&gt;Go back to an earlier state.&lt;/p&gt;

&lt;p&gt;Change something.&lt;/p&gt;

&lt;p&gt;Then continue from there.&lt;/p&gt;

&lt;p&gt;That idea led me to build an open-source &lt;strong&gt;time-travel debugger for AI agents&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;GitHub: &lt;a href="https://github.com/UjwalBagalkoti/ai-time-travel-debugger" rel="noopener noreferrer"&gt;https://github.com/UjwalBagalkoti/ai-time-travel-debugger&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  The Problem
&lt;/h2&gt;

&lt;p&gt;AI agents are becoming more capable, but debugging them is still surprisingly difficult.&lt;/p&gt;

&lt;p&gt;A typical agent might execute something like:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;User request
    ↓
LLM
    ↓
Search / Tool
    ↓
LLM
    ↓
Database / Tool
    ↓
LLM
    ↓
External API
    ↓
Final response
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Now imagine the final answer is wrong.&lt;/p&gt;

&lt;p&gt;The actual mistake may have happened several steps earlier.&lt;/p&gt;

&lt;p&gt;Maybe the model selected the wrong tool.&lt;/p&gt;

&lt;p&gt;Maybe a tool returned an unexpected result.&lt;/p&gt;

&lt;p&gt;Maybe the agent state changed unexpectedly.&lt;/p&gt;

&lt;p&gt;Maybe the model made a bad decision because of information introduced earlier in the execution.&lt;/p&gt;

&lt;p&gt;With traditional logging, you can inspect what happened.&lt;/p&gt;

&lt;p&gt;But if you want to experiment with what would have happened after changing an earlier decision, things become much harder.&lt;/p&gt;

&lt;p&gt;You often have to run the agent again from the beginning.&lt;/p&gt;

&lt;p&gt;That can mean:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;repeating model calls&lt;/li&gt;
&lt;li&gt;repeating tool calls&lt;/li&gt;
&lt;li&gt;making network requests again&lt;/li&gt;
&lt;li&gt;waiting for the entire workflow&lt;/li&gt;
&lt;li&gt;potentially triggering real-world side effects&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;I wanted a different approach.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Idea: Treat an Agent Execution as History
&lt;/h2&gt;

&lt;p&gt;Instead of thinking about an agent run as a stream of logs, I wanted to treat it as a historical execution that could be inspected.&lt;/p&gt;

&lt;p&gt;The workflow becomes:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Record
   ↓
Inspect
   ↓
Rewind
   ↓
Modify
   ↓
Replay
   ↓
Branch
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The important part is that the original execution remains available.&lt;/p&gt;

&lt;p&gt;You can use it as the starting point for another experiment.&lt;/p&gt;

&lt;p&gt;That's where the idea of &lt;strong&gt;time-travel debugging&lt;/strong&gt; comes from.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Architecture
&lt;/h2&gt;

&lt;p&gt;The current implementation is built around a relatively simple pipeline:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Python AgentTracer
        ↓
JSON execution trace
        ↓
FastAPI replay API
        ↓
PostgreSQL / SQLite
        ↓
Next.js + React Flow
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The project has four main parts.&lt;/p&gt;

&lt;h3&gt;
  
  
  1. Python tracing SDK
&lt;/h3&gt;

&lt;p&gt;The SDK records agent execution steps.&lt;/p&gt;

&lt;h3&gt;
  
  
  2. FastAPI backend
&lt;/h3&gt;

&lt;p&gt;The backend receives, stores, retrieves, and replays traces.&lt;/p&gt;

&lt;h3&gt;
  
  
  3. Database
&lt;/h3&gt;

&lt;p&gt;PostgreSQL is used in the deployed environment, while SQLite can be used for local development.&lt;/p&gt;

&lt;h3&gt;
  
  
  4. React Flow debugger
&lt;/h3&gt;

&lt;p&gt;The frontend visualizes the execution as an interactive graph and timeline.&lt;/p&gt;

&lt;p&gt;The goal was to keep each part relatively independent so that the debugger could eventually work with different agent frameworks.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Debugger Interface
&lt;/h2&gt;

&lt;p&gt;Once an execution is recorded, the frontend turns the trace into a visual execution graph.&lt;/p&gt;

&lt;p&gt;Instead of reading a large stream of logs, you can see the individual execution steps and select them for inspection.&lt;/p&gt;

&lt;p&gt;For example:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;LLM
 ↓
lookup_order
 ↓
LLM
 ↓
issue_refund
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Each step can be inspected through the debugger.&lt;/p&gt;

&lt;p&gt;The timeline scrubber also makes it possible to move through the recorded execution history.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F07yldtg79vxu0thsdcmt.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F07yldtg79vxu0thsdcmt.png" alt=" " width="800" height="453"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Caption: The AI Time-Travel Debugger showing an agent execution graph, timeline scrubber, step inspector, and execution metrics.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Alt text: AI agent time-travel debugger displaying a four-step execution graph with an LLM call, tool execution, second LLM call, and final tool execution.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;This gives me a much clearer view of the execution than a traditional log file.&lt;/p&gt;

&lt;p&gt;I can see the trajectory, select an individual step, and inspect the recorded information associated with it.&lt;/p&gt;

&lt;h2&gt;
  
  
  Recording an Agent Execution
&lt;/h2&gt;

&lt;p&gt;The first requirement was capturing enough information about an execution to make it useful later.&lt;/p&gt;

&lt;p&gt;The Python tracer records information such as:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;step type&lt;/li&gt;
&lt;li&gt;arguments&lt;/li&gt;
&lt;li&gt;outputs&lt;/li&gt;
&lt;li&gt;state snapshots&lt;/li&gt;
&lt;li&gt;status&lt;/li&gt;
&lt;li&gt;timestamps&lt;/li&gt;
&lt;li&gt;latency&lt;/li&gt;
&lt;li&gt;model information&lt;/li&gt;
&lt;li&gt;tool calls&lt;/li&gt;
&lt;li&gt;token metrics where available&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;For example, a recorded tool execution can contain information like:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"tool_name"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"lookup_order"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"arguments"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"order_id"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"992"&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;and its recorded result:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"result"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"status"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"delivered"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"amount"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;120&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"refundable"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="kc"&gt;true&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The important thing isn't simply collecting more logs.&lt;/p&gt;

&lt;p&gt;The goal is to preserve enough execution history to investigate a failure.&lt;/p&gt;

&lt;h2&gt;
  
  
  Inspecting Individual Steps
&lt;/h2&gt;

&lt;p&gt;One of the useful parts of the debugger is being able to select an individual step.&lt;/p&gt;

&lt;p&gt;For example, selecting the &lt;code&gt;lookup_order&lt;/code&gt; step shows the recorded tool arguments and the output that was produced during the original execution.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F6hrfl7w3wycq5xd5r5w1.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F6hrfl7w3wycq5xd5r5w1.png" alt=" " width="800" height="419"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Caption: Inspecting a recorded &lt;code&gt;lookup_order&lt;/code&gt; tool execution and its recorded output.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Alt text: Debugger inspector showing the lookup_order tool, order ID 992, and the recorded result showing the order as delivered and refundable.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;This is important for replay because the debugger has a historical record of what the tool returned.&lt;/p&gt;

&lt;p&gt;Instead of treating the entire execution as one opaque operation, each step can be examined separately.&lt;/p&gt;

&lt;h2&gt;
  
  
  Deterministic Tool Replay
&lt;/h2&gt;

&lt;p&gt;This was one of the most important parts of the project.&lt;/p&gt;

&lt;p&gt;Consider a tool call:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;lookup_order(order_id="992")
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;During the original execution, the tool returned:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"status"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"delivered"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"amount"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;120&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"refundable"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="kc"&gt;true&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;If we replay the same execution and encounter the same tool with the same arguments, we don't necessarily need to execute the external tool again.&lt;/p&gt;

&lt;p&gt;Instead, the debugger can use the recorded result.&lt;/p&gt;

&lt;p&gt;Conceptually:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Original execution

lookup_order("992")
        ↓
external tool
        ↓
record result
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Later:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Replay

lookup_order("992")
        ↓
recorded result
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This is useful for both speed and safety.&lt;/p&gt;

&lt;p&gt;It also means that debugging doesn't automatically require another network request.&lt;/p&gt;

&lt;h2&gt;
  
  
  Inspecting the LLM Decision
&lt;/h2&gt;

&lt;p&gt;The tool result is only one part of an agent execution.&lt;/p&gt;

&lt;p&gt;The next step may be an LLM decision based on that information.&lt;/p&gt;

&lt;p&gt;In the example execution, the next LLM step contains the instruction:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Authorize refund based on order details.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The debugger allows this intermediate step and its recorded output to be inspected.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fyjrzko5c040mo9dn6bli.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fyjrzko5c040mo9dn6bli.png" alt=" " width="799" height="429"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Caption: Inspecting the intermediate LLM step that determines whether the order is eligible for a refund.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Alt text: AI agent debugger showing step 3 as an LLM call with the prompt "Authorize refund based on order details" and its recorded output.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;This is where the historical execution becomes useful for debugging.&lt;/p&gt;

&lt;p&gt;Instead of only seeing the final result, I can inspect the intermediate point where the agent made its decision.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why Replay Needs a Safety Boundary
&lt;/h2&gt;

&lt;p&gt;At first, replay sounds simple.&lt;/p&gt;

&lt;p&gt;Just execute everything again.&lt;/p&gt;

&lt;p&gt;But that's exactly where things can go wrong.&lt;/p&gt;

&lt;p&gt;Imagine an agent has tools such as:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;send_email()
charge_card()
delete_database_record()
issue_refund()
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;If a debugger blindly re-executed every tool during replay, debugging could accidentally cause real-world side effects.&lt;/p&gt;

&lt;p&gt;A debugging tool should not surprise you by sending an email or issuing a refund.&lt;/p&gt;

&lt;p&gt;So I designed the replay system around a &lt;strong&gt;safe replay boundary&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;Recorded tool calls can be replayed when their tool name and arguments match the recorded execution.&lt;/p&gt;

&lt;p&gt;Unknown external calls aren't silently executed.&lt;/p&gt;

&lt;p&gt;Instead, replay can stop at the boundary.&lt;/p&gt;

&lt;p&gt;The principle is:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Replay what is known. Don't blindly execute what isn't.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;This is still an area with many problems to solve, but I think making replay safe by default is an important foundation.&lt;/p&gt;

&lt;h2&gt;
  
  
  Branching an Execution
&lt;/h2&gt;

&lt;p&gt;The next idea was branching.&lt;/p&gt;

&lt;p&gt;Suppose an agent reaches this point:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;1 → 2 → 3 → 4 → 5 → 6
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;At step 4, the agent makes a decision that leads to the wrong outcome.&lt;/p&gt;

&lt;p&gt;With a traditional workflow, you might restart everything.&lt;/p&gt;

&lt;p&gt;With a time-travel debugger, the goal is to create another trajectory:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;             ┌→ 5 → 6
1 → 2 → 3 → 4
             └→ 4' → 5' → 6'
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The original execution remains untouched.&lt;/p&gt;

&lt;p&gt;The new execution becomes a separate branch.&lt;/p&gt;

&lt;p&gt;The debugger exposes this through the &lt;strong&gt;Fork &amp;amp; Replay Branch&lt;/strong&gt; workflow.&lt;/p&gt;

&lt;p&gt;The historical execution effectively becomes a starting point for another experiment.&lt;/p&gt;

&lt;h2&gt;
  
  
  What Makes This Different From Logs?
&lt;/h2&gt;

&lt;p&gt;Logs are extremely useful.&lt;/p&gt;

&lt;p&gt;But logs generally answer:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;What happened?&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;A replay-oriented debugger aims to help answer another question:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;What could happen if I change something that happened earlier?&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;That's the fundamental difference I'm exploring.&lt;/p&gt;

&lt;p&gt;A normal log might tell you:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Step 1 completed
Step 2 completed
Step 3 failed
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;A time-travel workflow tries to give you:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Step 1
  ↓
Step 2
  ↓
rewind
  ↓
modify
  ↓
replay
  ↓
new execution branch
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;For increasingly complex agent systems, that distinction could become important.&lt;/p&gt;

&lt;h2&gt;
  
  
  A Small Example
&lt;/h2&gt;

&lt;p&gt;Imagine an agent responsible for handling an order.&lt;/p&gt;

&lt;p&gt;The execution is:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;User request
 ↓
LLM
 ↓
lookup_order
 ↓
LLM
 ↓
issue_refund
 ↓
Final response
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;In my example trace, the agent looks up order &lt;code&gt;#992&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;The tool returns:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Status: delivered
Amount: $120
Refundable: true
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The next LLM step evaluates the order details, and the final tool execution records the refund action.&lt;/p&gt;

&lt;p&gt;The debugger lets each of these steps be inspected independently.&lt;/p&gt;

&lt;p&gt;The goal is not just to see the final answer.&lt;/p&gt;

&lt;p&gt;The goal is to understand the path that produced it.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Current Stack
&lt;/h2&gt;

&lt;p&gt;The project currently uses:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Python&lt;/li&gt;
&lt;li&gt;FastAPI&lt;/li&gt;
&lt;li&gt;PostgreSQL&lt;/li&gt;
&lt;li&gt;SQLite for local development&lt;/li&gt;
&lt;li&gt;Next.js&lt;/li&gt;
&lt;li&gt;React&lt;/li&gt;
&lt;li&gt;React Flow&lt;/li&gt;
&lt;li&gt;Docker&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The repository also contains a Python tracing SDK, replay engine tests, production end-to-end tests, and an example trace.&lt;/p&gt;

&lt;h2&gt;
  
  
  What I Haven't Solved Yet
&lt;/h2&gt;

&lt;p&gt;This project is still an early implementation.&lt;/p&gt;

&lt;p&gt;There are many hard problems remaining.&lt;/p&gt;

&lt;h3&gt;
  
  
  Deterministic LLM Replay
&lt;/h3&gt;

&lt;p&gt;LLMs aren't simple deterministic functions.&lt;/p&gt;

&lt;p&gt;Reproducing an exact model trajectory can be difficult depending on the model, provider, configuration, tools, and surrounding state.&lt;/p&gt;

&lt;h3&gt;
  
  
  Nondeterministic Tools
&lt;/h3&gt;

&lt;p&gt;Some tools depend on external state.&lt;/p&gt;

&lt;p&gt;For example:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;current_weather()
stock_price()
web_search()
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Their results can change between executions.&lt;/p&gt;

&lt;p&gt;Caching their outputs helps replay, but it doesn't solve every reproducibility problem.&lt;/p&gt;

&lt;h3&gt;
  
  
  Side Effects
&lt;/h3&gt;

&lt;p&gt;A real agent can modify the outside world.&lt;/p&gt;

&lt;p&gt;That creates a fundamental question:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;How should a debugger safely replay an operation that changes something outside the agent?&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h3&gt;
  
  
  Large State
&lt;/h3&gt;

&lt;p&gt;Agent state can become large.&lt;/p&gt;

&lt;p&gt;Recording every state snapshot may eventually become expensive.&lt;/p&gt;

&lt;p&gt;A production system needs strategies for:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;compression&lt;/li&gt;
&lt;li&gt;deduplication&lt;/li&gt;
&lt;li&gt;incremental snapshots&lt;/li&gt;
&lt;li&gt;storage lifecycle&lt;/li&gt;
&lt;li&gt;selective recording&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  Concurrency
&lt;/h3&gt;

&lt;p&gt;Real agents may execute multiple operations concurrently.&lt;/p&gt;

&lt;p&gt;A simple linear timeline isn't enough to represent every possible execution.&lt;/p&gt;

&lt;h3&gt;
  
  
  Streaming
&lt;/h3&gt;

&lt;p&gt;LLM responses and tool outputs can also be streamed.&lt;/p&gt;

&lt;p&gt;Capturing and replaying those partial events introduces another layer of complexity.&lt;/p&gt;

&lt;h3&gt;
  
  
  Framework Integrations
&lt;/h3&gt;

&lt;p&gt;Different agent frameworks have different execution models.&lt;/p&gt;

&lt;p&gt;A useful debugger eventually needs to integrate naturally with the ecosystems developers already use.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where I Want to Take It
&lt;/h2&gt;

&lt;p&gt;The current implementation is a foundation.&lt;/p&gt;

&lt;p&gt;The larger direction I'm interested in is making agent execution something developers can inspect, reproduce, and experiment with.&lt;/p&gt;

&lt;p&gt;That could eventually mean deeper integrations with:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;agent frameworks&lt;/li&gt;
&lt;li&gt;observability systems&lt;/li&gt;
&lt;li&gt;tool execution runtimes&lt;/li&gt;
&lt;li&gt;distributed workflows&lt;/li&gt;
&lt;li&gt;evaluation systems&lt;/li&gt;
&lt;li&gt;production debugging infrastructure&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;There are still many architectural questions I don't have answers to.&lt;/p&gt;

&lt;p&gt;And that's part of why I open-sourced it.&lt;/p&gt;

&lt;h2&gt;
  
  
  Open Source
&lt;/h2&gt;

&lt;p&gt;The project is available on GitHub:&lt;/p&gt;

&lt;p&gt;&lt;a href="https://github.com/UjwalBagalkoti/ai-time-travel-debugger" rel="noopener noreferrer"&gt;https://github.com/UjwalBagalkoti/ai-time-travel-debugger&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The repository contains the backend, frontend, SDK, replay engine, tests, Docker configuration, and example trace.&lt;/p&gt;

&lt;p&gt;The deployed application is also available to experiment with:&lt;/p&gt;

&lt;p&gt;&lt;a href="https://ai-time-travel-debugger.onrender.com" rel="noopener noreferrer"&gt;https://ai-time-travel-debugger.onrender.com&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;I'm especially interested in feedback from developers building:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;AI agents&lt;/li&gt;
&lt;li&gt;tool-calling systems&lt;/li&gt;
&lt;li&gt;agent runtimes&lt;/li&gt;
&lt;li&gt;LangChain/LangGraph applications&lt;/li&gt;
&lt;li&gt;Python agent systems&lt;/li&gt;
&lt;li&gt;observability infrastructure&lt;/li&gt;
&lt;li&gt;developer debugging tools&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The question I'm most interested in is:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;What information must an agent runtime capture so that a failed execution can actually be reproduced and investigated?&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;If you're building agents, I'd love to hear how you're currently debugging failures.&lt;/p&gt;

&lt;p&gt;What do you record?&lt;/p&gt;

&lt;p&gt;What do you wish you had recorded?&lt;/p&gt;

&lt;p&gt;And where does replay break down in your systems?&lt;/p&gt;

&lt;h2&gt;
  
  
  Final Thought
&lt;/h2&gt;

&lt;p&gt;AI agents are starting to look less like simple functions and more like small distributed programs.&lt;/p&gt;

&lt;p&gt;They make decisions.&lt;/p&gt;

&lt;p&gt;They call tools.&lt;/p&gt;

&lt;p&gt;They maintain state.&lt;/p&gt;

&lt;p&gt;They interact with external systems.&lt;/p&gt;

&lt;p&gt;And when something goes wrong, "just run it again" isn't always enough.&lt;/p&gt;

&lt;p&gt;That's the problem I'm exploring with this project.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Record once.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Inspect the past.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Rewind the execution.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Experiment with another path.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Replay safely.&lt;/strong&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>python</category>
      <category>opensource</category>
      <category>devtools</category>
    </item>
  </channel>
</rss>
