DEV Community

Cover image for Why Code Diffs Are Not Enough for AI Agent Changes
Raju Dandigam
Raju Dandigam

Posted on

Why Code Diffs Are Not Enough for AI Agent Changes

A pull request changes three lines in a prompt. The source diff is tiny. The resulting agent run adds a tool call, skips a planning step, changes its recovery path, and takes twice as long.

Which diff describes the risk more accurately?

Both do—but they describe different things.

I maintain AgentInspect, an open-source TypeScript toolkit for inspecting agent runs locally. I designed its run-diff workflow around a simple idea: code review tells us what the developer changed; execution evidence tells us what the agent did differently. The examples below use synthetic fixtures verified against agent-inspect@6.17.4.

Source changes and behavior are no longer tightly coupled

In ordinary deterministic code, a source diff is often a strong predictor of runtime change. Agent systems add several moving parts:

  • prompts and system instructions;
  • model versions and sampling behavior;
  • tool descriptions and schemas;
  • retrieval content;
  • external services;
  • memory and conversation state;
  • orchestration and fallback policies.

A large refactor can preserve the same trajectory. A one-word prompt change can alter tool selection. Identical code can behave differently when a model or external result changes.

That does not make code review obsolete. It means the review unit needs another layer:

source diff                    behavior diff
-----------                    -------------
what was edited?               what path changed?
what code owns the change?     where did runs diverge?
is the implementation sound?   which steps/errors/outputs differ?
Enter fullscreen mode Exit fullscreen mode

Capture two comparable runs

A useful behavioral comparison needs a deliberate baseline and candidate:

baseline
  same synthetic input
  pinned or recorded configuration
  known-good trace

candidate
  same synthetic input
  intended code/prompt/model change
  newly captured trace
Enter fullscreen mode Exit fullscreen mode

Control what you can. Record what you cannot. If the input, retrieval corpus, model, and tool fixtures all change at once, the diff may be accurate but difficult to interpret.

With two local run IDs, compare them from the CLI:

npx agent-inspect diff minimal-success minimal-error \
  --dir .agent-inspect
Enter fullscreen mode Exit fullscreen mode

The synthetic fixture output begins with the summary and first divergence:

Run diff
Left:  minimal-success
Right: minimal-error

Summary:
  Differences: 4
  Errors: 0
  Warnings: 3
  Info: 1

First divergence:
  run-status at (run)
    left: success
    right: error
Enter fullscreen mode Exit fullscreen mode

That “first divergence” is often more actionable than a long list of event differences. It gives the reviewer a starting point for the causal investigation.

Read added and removed steps as structural evidence

The same fixture reports:

Differences:
  [warning] run-status
    Run completion status differs
    left: success
    right: error
  [info] duration
    Run duration differs
    left: 120
    right: 70
  [warning] step-removed plan
    Step only in left run: plan
    left: step_root
    right: (undefined)
  [warning] step-added failing-step
    Step only in right run: failing-step
    left: (undefined)
    right: step_fail
Enter fullscreen mode Exit fullscreen mode

The evidence says that plan appears only in the left run and failing-step only in the right. It does not automatically say why.

Possible explanations include:

  • the candidate genuinely skipped planning;
  • a step was renamed;
  • instrumentation boundaries changed;
  • the agent chose a different path;
  • the run ended before reaching the step.

This is why a behavioral diff is an input to review, not an automatic verdict. Pair it with the source diff and inspect the execution tree around the divergence.

Focus the diff on the question you are asking

The CLI can limit the comparison to a specific check dimension. To inspect only structure:

npx agent-inspect diff minimal-success minimal-error \
  --dir .agent-inspect \
  --check structure
Enter fullscreen mode Exit fullscreen mode

That removes status and duration noise and leaves the added/removed steps. For a performance-oriented review:

npx agent-inspect diff minimal-success minimal-error \
  --dir .agent-inspect \
  --check timing \
  --duration-threshold 20ms
Enter fullscreen mode Exit fullscreen mode

The command also supports JSON output for automation, --ignore-duration, focus modes, and verbose output. A practical rule is to begin with the broad human-readable diff, then narrow the view when you know which hypothesis you are testing.

Timing differences require controlled interpretation

The fixture reports 120 versus 70 milliseconds. It would be a mistake to generalize that single synthetic delta into a performance claim.

Agent latency can vary with network conditions, cache state, provider load, token volume, and concurrency. A timing diff is most useful when:

  • the tool and model calls are stubbed in a deterministic test;
  • the difference is large relative to expected noise;
  • multiple representative runs show the same pattern;
  • the structural diff explains the extra work.

Use --duration-threshold to suppress insignificant changes, but derive the threshold from your environment. Do not choose a number merely because it makes a current test pass.

A run diff does not rerun the agent

AgentInspect’s diff is a read-only comparison of persisted traces. It does not replay either agent, invoke a model, or prove that the difference will recur.

That property is useful for review: the comparison is deterministic for the two stored artifacts. It is also a limitation: representative capture remains your responsibility.

I think of the workflow as three separate actions:

execute -> capture evidence
compare -> describe observed differences
judge   -> decide whether the change is acceptable
Enter fullscreen mode Exit fullscreen mode

Only the middle action is the run-diff engine.

Add a behavior-evidence section to pull requests

For changes that can materially affect an agent path, a compact pull-request section can make review faster:

## Agent behavior evidence

- Fixture: `refund-eligible-order`
- Baseline run: `refund-before`
- Candidate run: `refund-after`
- Expected change: prefer cached policy when current
- First divergence: `retrieve-policy` replaced by `load-policy-cache`
- Contract result: pass
- Evidence artifact: attached CI bundle
- Reviewer note: no production or customer data used
Enter fullscreen mode Exit fullscreen mode

This is intentionally concise. The full trace should remain an artifact, not be pasted into the pull-request description.

The most important field is “expected change.” It tells reviewers whether an observed divergence is intentional. Without that statement, the diff is merely a list of facts.

Combine diffing with contracts

A diff answers “What changed?” A contract answers “Did a declared invariant still hold?” Use both.

Suppose a candidate run replaces remote retrieval with a cache hit. The structural diff should show the path change. A contract might still require:

  • successful completion;
  • no forbidden write tool;
  • validation before the final response;
  • a bounded tool-call count.

The candidate can therefore differ from the baseline and still pass the invariant gate. This is healthier than either extreme:

  • rejecting every structural change; or
  • accepting every change that ends with plausible prose.

What to compare in practice

Choose fixtures around decisions, not around random traffic. High-value comparisons include:

Prompt or instruction changes

Did tool selection, ordering, retry behavior, or token usage change?

Model upgrades

Does the new model reach the same goal with a different trajectory? Does it invoke a fallback more often in controlled cases?

Tool-schema changes

Did a renamed or re-described tool disappear, get replaced, or start failing?

Orchestrator refactors

Did parent-child structure, concurrency, or error propagation change even if final answers stayed stable?

Retrieval changes

Did the run add or skip retrieval, or generate before the required evidence step?

For each case, pair the behavioral observation with a domain-specific quality check. A shorter path is not necessarily a better answer.

Review what executed, not only what was edited

Code diffs remain the foundation of software review. Agent systems need an additional artifact because runtime behavior depends on more than source text.

A disciplined workflow is straightforward:

  1. capture a known-good baseline for a synthetic, representative fixture;
  2. capture the candidate with controlled inputs;
  3. inspect the first divergence and focused structural changes;
  4. run stable deterministic contracts;
  5. preserve the evidence with the pull request;
  6. apply human and semantic judgment to the result.

This does not eliminate nondeterminism. It gives reviewers something more precise than “I tried the prompt and it looked better.”

The tagged release and fixtures used for this article are available on GitHub. If you adopt behavior diffs in code review, begin with one high-risk fixture and learn which differences your team actually finds actionable.

Top comments (18)

Collapse
 
kevinpruett023_kevinpruet profile image
Lee •

This is also good. I wanna discuss further about collaboration via telegram.
@morganruiz5363 this is my tg. how about you?

Collapse
 
raju_dandigam profile image
Raju Dandigam •

@kevinpruett023_kevinpruet, I keep technical collaboration on DEV or GitHub rather than Telegram. If you have a concrete question about behavior diffs—fixture selection, first-divergence semantics, or CI evidence bundles—please add it here so the context stays reviewable for everyone.

Collapse
 
kevinpruett023_kevinpruet profile image
Lee •

Okay.

Thread Thread
 
raju_dandigam profile image
Raju Dandigam •

@kevinpruett023_kevinpruet, thanks for understanding. If you have a concrete behavior-diff case, add it here and I’ll be glad to discuss the fixture, first divergence, or evidence contract with you.

Thread Thread
 
kevinpruett023_kevinpruet profile image
Lee •

Thanks. I can share a behavior-diff case around an API regression where two implementations produced different outputs under the same input contract.

The fixture was a set of production-like requests containing:

-nested JSON payloads
-optional fields
-edge-case null values
-ordering variations
-invalid-but-tolerated inputs

The first divergence appeared in the response normalization layer:

Expected behavior:

-Preserve missing fields as null
-Maintain deterministic ordering
-Return the same validation error schema

Observed behavior:

-One implementation dropped optional fields
-Error responses changed from structured validation objects to generic exceptions
-Some fields were reordered, breaking downstream snapshot comparisons

My approach would be:

1.Freeze the failing fixture as a regression test.
2.Compare execution traces between both versions.
3.Identify the first semantic divergence rather than the final visible failure.
4.Define the evidence contract:

  • input fixture
  • expected output
  • allowed variations
  • invariants that must remain unchanged

For debugging, I usually focus on:

-request transformation boundaries
-serialization/deserialization behavior
-dependency version changes
-hidden state/cache effects
-asynchronous execution ordering

The key is not only proving that outputs differ, but proving where and why the behavior contract changed.

A more AI-focused version (if the discussion is about LLM/AI systems):

I have seen similar behavior-diff issues in RAG systems where the same query produced different answers after changing embedding models or retrieval parameters.

The fixture included:

-fixed documents
-fixed query set
-expected retrieved chunks
-answer evaluation criteria

The first divergence was not the final answer quality; it was the retrieval stage where chunk ranking changed. That caused downstream generation differences.

The evidence contract included:

-embedding version
-chunking strategy
-top-k retrieval results
-prompt template
-model parameters
-final answer evaluation score

This allowed us to isolate whether the regression came from retrieval, prompting, or generation.

I look forward to more active discussion in the future.
Best

Thread Thread
 
raju_dandigam profile image
Raju Dandigam •

@kevinpruett023_kevinpruet, this is a useful concrete fixture. I’d normalize ordering only where order is semantically irrelevant, but keep null-versus-missing and the validation-error envelope as explicit invariants; otherwise canonicalization can hide the regression. For the RAG case, recording ranked chunk IDs and scores would let the first divergence be attributed to retrieval before judging the final answer. Thanks for laying out both versions of the evidence contract.

Thread Thread
 
kevinpruett023_kevinpruet profile image
Lee •

Thanks for the thoughtful feedback. I agree that canonicalization needs to be applied carefully. The goal is not to normalize away meaningful differences, but to separate representation-level noise from actual contract violations.

For API behavior diffs, I usually classify fields into three categories:

  • strict invariants: values where any change represents a regression (for example, validation schemas, required fields, security-related metadata)
  • nullable/optional semantics: where missing, null, and default values must be explicitly defined
  • non-semantic fields: where ordering or formatting differences can be normalized safely

For RAG systems, I completely agree that retrieval-level observability is critical. In practice, the final answer is often only a symptom. Capturing ranked chunk IDs, similarity scores, embedding model versions, chunking configuration, and retrieval parameters creates a traceable chain from query → retrieval → context → generation.

One additional layer I usually add is automated evaluation at each stage:

  • retrieval metrics (precision@k, recall@k, relevance scoring)
  • context quality checks
  • groundedness and hallucination evaluation
  • final response quality scoring

This makes it possible to identify whether a regression comes from data changes, embedding drift, retrieval configuration, prompt changes, or model behavior.

The main principle is that regression debugging should follow the earliest point where the system behavior diverges, not the point where the user-visible failure appears.

If there are any collaborative projects, I would like to work on you together.

Best

Thread Thread
 
raju_dandigam profile image
Raju Dandigam •

@kevinpruett023_kevinpruet, that stage-by-stage evaluation split is the right shape, especially pinning the embedding and chunking configuration beside ranked results. I’d keep retrieval metrics and answer-quality metrics as separate evidence lanes: stable recall@k can still hide different chunk identities or ordering, while a retrieval regression can explain a downstream answer change. A useful collaboration artifact could be a small synthetic RAG fixture with a trace from query → ranked chunks → assembled context → answer, deterministic invariants before semantic grading. If you’re interested, that would be worth developing publicly on DEV or GitHub.

Thread Thread
 
kevinpruett023_kevinpruet profile image
Lee •

okay. no problem

Thread Thread
 
raju_dandigam profile image
Raju Dandigam •

@kevinpruett023_kevinpruet, thanks. If you decide to prototype that public RAG fixture, a focused first version could freeze one query set and assert ranked chunk IDs, context assembly, and a few deterministic invariants before semantic scoring. Feel free to share the artifact here or on GitHub when you’re ready.

Thread Thread
 
kevinpruett023_kevinpruet profile image
Lee •

Okay...No problem. Best

Thread Thread
 
raju_dandigam profile image
Raju Dandigam •

@kevinpruett023_kevinpruet, thanks. If you put the fixture together, share the frozen query set and the retrieval invariants you want to enforce; I’d be glad to compare the trace design with you here.

Collapse
 
vinhnguyenthanhdn profile image
Vinh Nguyen •

Your list of explanations for a step-removed entry has one item that is not a peer of the others: "the run ended before reaching the step" also manufactures the symptoms of every other item on the list. A candidate that errors early reports each step after the failure as removed, so Differences: 4 scales with how early the run stopped rather than with how much behaviour changed — a candidate failing at step 2 of 20 reads as a large divergence and one failing at step 19 reads as a small one, from the same defect. It also puts "first divergence" on the truncation boundary instead of on the cause, which is the opposite of what you want it for.

The shape of the removed set separates those cheaply, with no extra capture. If the steps present only in the baseline form a contiguous suffix of the baseline trajectory, you are reading truncation and the structural diff carries close to no behavioural information. If they are interior — something missing with steps still following it — the path genuinely changed. Your own fixture is the ambiguous case: plan removed and failing-step added is consistent both with a candidate that skipped planning and with a candidate that died before it, and only the position of plan in the baseline tells you which.

--check structure makes this sharper rather than safer, which is worth a line in the docs. It strips run-status, and run-status is the field carrying right: error — the only signal that the removed steps might be a truncation artifact. So the mode advertised as removing status and duration noise removes exactly the evidence needed to interpret what it leaves behind.

Collapse
 
raju_dandigam profile image
Raju Dandigam •

@vinhnguyenthanhdn, you’re right: truncation should not be counted as N independent structural divergences. A contiguous baseline-only suffix after a candidate terminal error is better classified as not_reached, while an interior absence is a genuine path change; the terminal status must remain attached even in a structure-only view. In this fixture, downstream omissions should be attributed to termination rather than reported as separate removals. I’ll correct the explanation and inspect the CLI behavior so --check structure never discards the evidence needed to classify that shape.

Collapse
 
salparvez profile image
Sal Parvez | ML Systems •

The "Agent behavior evidence" block is a record with three different kinds of entry in it, and it reads better once they're labeled. The first divergence and the added/removed steps are measured. The expected change is stated, by the developer, before the run. The contract result is a verification. Review is the act of checking that the stated line accounts for the measured ones, and it fails in a specific way when the expected-change line is written after looking at the diff, because then it is a description, not a prediction.

The addition I'd make to the PR section: bind the reviewer's approval to a hash of the evidence bundle. If the bundle is regenerated after approval, the approval should lapse rather than carry forward, or the section becomes a record of what someone once looked at.

Collapse
 
raju_dandigam profile image
Raju Dandigam •

@salparvez, the measured/stated/verified separation is much sharper. “Expected change” should be committed before the candidate evidence exists—or at least anchored to the relevant PR revision—so it remains a prediction rather than a description written with hindsight. Binding approval to a hash of the evidence bundle, together with the baseline and candidate commits, fixture, and configuration, would make the reviewed object immutable. Regenerating any part would produce a new hash and require fresh approval. I’ll add those labels and the evidence digest to the proposed PR block.

Collapse
 
salparvez profile image
Sal Parvez | ML Systems •

Committing the expected change before the candidate evidence exists is the right cut. It is the same reason a stamp has to be bound to a content hash: the reviewed object has to be immutable, or the review is a description written with hindsight. Baseline commit, candidate commit, fixture, configuration, digest of the bundle, one approval over all of it, and a regenerated part is a new object needing a new approval. Looking forward to the revised block.

Thread Thread
 
raju_dandigam profile image
Raju Dandigam •

@salparvez, agreed. I’m going to treat that tuple as the review identity rather than metadata around the review: baseline commit, candidate commit, fixture and configuration digest, and Evidence bundle digest. The PR block should show the expected change committed against that identity, and any regeneration should invalidate approval instead of silently refreshing the artifact. I also think the verifier should fail if the displayed tuple and manifest do not match, so the rule is enforced rather than only documented.