DEV Community

Cover image for Why Code Diffs Are Not Enough for AI Agent Changes

Why Code Diffs Are Not Enough for AI Agent Changes

Raju Dandigam on September 07, 2026

A pull request changes three lines in a prompt. The source diff is tiny. The resulting agent run adds a tool call, skips a planning step, changes i...
Collapse
 
kevinpruett023_kevinpruet profile image
Lee •

This is also good. I wanna discuss further about collaboration via telegram.
@morganruiz5363 this is my tg. how about you?

Collapse
 
raju_dandigam profile image
Raju Dandigam •

@kevinpruett023_kevinpruet, I keep technical collaboration on DEV or GitHub rather than Telegram. If you have a concrete question about behavior diffs—fixture selection, first-divergence semantics, or CI evidence bundles—please add it here so the context stays reviewable for everyone.

Collapse
 
kevinpruett023_kevinpruet profile image
Lee •

Okay.

Thread Thread
 
raju_dandigam profile image
Raju Dandigam •

@kevinpruett023_kevinpruet, thanks for understanding. If you have a concrete behavior-diff case, add it here and I’ll be glad to discuss the fixture, first divergence, or evidence contract with you.

Thread Thread
 
kevinpruett023_kevinpruet profile image
Lee •

Thanks. I can share a behavior-diff case around an API regression where two implementations produced different outputs under the same input contract.

The fixture was a set of production-like requests containing:

-nested JSON payloads
-optional fields
-edge-case null values
-ordering variations
-invalid-but-tolerated inputs

The first divergence appeared in the response normalization layer:

Expected behavior:

-Preserve missing fields as null
-Maintain deterministic ordering
-Return the same validation error schema

Observed behavior:

-One implementation dropped optional fields
-Error responses changed from structured validation objects to generic exceptions
-Some fields were reordered, breaking downstream snapshot comparisons

My approach would be:

1.Freeze the failing fixture as a regression test.
2.Compare execution traces between both versions.
3.Identify the first semantic divergence rather than the final visible failure.
4.Define the evidence contract:

  • input fixture
  • expected output
  • allowed variations
  • invariants that must remain unchanged

For debugging, I usually focus on:

-request transformation boundaries
-serialization/deserialization behavior
-dependency version changes
-hidden state/cache effects
-asynchronous execution ordering

The key is not only proving that outputs differ, but proving where and why the behavior contract changed.

A more AI-focused version (if the discussion is about LLM/AI systems):

I have seen similar behavior-diff issues in RAG systems where the same query produced different answers after changing embedding models or retrieval parameters.

The fixture included:

-fixed documents
-fixed query set
-expected retrieved chunks
-answer evaluation criteria

The first divergence was not the final answer quality; it was the retrieval stage where chunk ranking changed. That caused downstream generation differences.

The evidence contract included:

-embedding version
-chunking strategy
-top-k retrieval results
-prompt template
-model parameters
-final answer evaluation score

This allowed us to isolate whether the regression came from retrieval, prompting, or generation.

I look forward to more active discussion in the future.
Best

Thread Thread
 
raju_dandigam profile image
Raju Dandigam •

@kevinpruett023_kevinpruet, this is a useful concrete fixture. I’d normalize ordering only where order is semantically irrelevant, but keep null-versus-missing and the validation-error envelope as explicit invariants; otherwise canonicalization can hide the regression. For the RAG case, recording ranked chunk IDs and scores would let the first divergence be attributed to retrieval before judging the final answer. Thanks for laying out both versions of the evidence contract.

Collapse
 
vinhnguyenthanhdn profile image
Vinh Nguyen •

Your list of explanations for a step-removed entry has one item that is not a peer of the others: "the run ended before reaching the step" also manufactures the symptoms of every other item on the list. A candidate that errors early reports each step after the failure as removed, so Differences: 4 scales with how early the run stopped rather than with how much behaviour changed — a candidate failing at step 2 of 20 reads as a large divergence and one failing at step 19 reads as a small one, from the same defect. It also puts "first divergence" on the truncation boundary instead of on the cause, which is the opposite of what you want it for.

The shape of the removed set separates those cheaply, with no extra capture. If the steps present only in the baseline form a contiguous suffix of the baseline trajectory, you are reading truncation and the structural diff carries close to no behavioural information. If they are interior — something missing with steps still following it — the path genuinely changed. Your own fixture is the ambiguous case: plan removed and failing-step added is consistent both with a candidate that skipped planning and with a candidate that died before it, and only the position of plan in the baseline tells you which.

--check structure makes this sharper rather than safer, which is worth a line in the docs. It strips run-status, and run-status is the field carrying right: error — the only signal that the removed steps might be a truncation artifact. So the mode advertised as removing status and duration noise removes exactly the evidence needed to interpret what it leaves behind.

Collapse
 
raju_dandigam profile image
Raju Dandigam •

@vinhnguyenthanhdn, you’re right: truncation should not be counted as N independent structural divergences. A contiguous baseline-only suffix after a candidate terminal error is better classified as not_reached, while an interior absence is a genuine path change; the terminal status must remain attached even in a structure-only view. In this fixture, downstream omissions should be attributed to termination rather than reported as separate removals. I’ll correct the explanation and inspect the CLI behavior so --check structure never discards the evidence needed to classify that shape.

Collapse
 
salparvez profile image
Sal Parvez | ML Systems •

The "Agent behavior evidence" block is a record with three different kinds of entry in it, and it reads better once they're labeled. The first divergence and the added/removed steps are measured. The expected change is stated, by the developer, before the run. The contract result is a verification. Review is the act of checking that the stated line accounts for the measured ones, and it fails in a specific way when the expected-change line is written after looking at the diff, because then it is a description, not a prediction.

The addition I'd make to the PR section: bind the reviewer's approval to a hash of the evidence bundle. If the bundle is regenerated after approval, the approval should lapse rather than carry forward, or the section becomes a record of what someone once looked at.

Collapse
 
raju_dandigam profile image
Raju Dandigam •

@salparvez, the measured/stated/verified separation is much sharper. “Expected change” should be committed before the candidate evidence exists—or at least anchored to the relevant PR revision—so it remains a prediction rather than a description written with hindsight. Binding approval to a hash of the evidence bundle, together with the baseline and candidate commits, fixture, and configuration, would make the reviewed object immutable. Regenerating any part would produce a new hash and require fresh approval. I’ll add those labels and the evidence digest to the proposed PR block.

Collapse
 
salparvez profile image
Sal Parvez | ML Systems •

Committing the expected change before the candidate evidence exists is the right cut. It is the same reason a stamp has to be bound to a content hash: the reviewed object has to be immutable, or the review is a description written with hindsight. Baseline commit, candidate commit, fixture, configuration, digest of the bundle, one approval over all of it, and a regenerated part is a new object needing a new approval. Looking forward to the revised block.

Thread Thread
 
raju_dandigam profile image
Raju Dandigam •

@salparvez, agreed. I’m going to treat that tuple as the review identity rather than metadata around the review: baseline commit, candidate commit, fixture and configuration digest, and Evidence bundle digest. The PR block should show the expected change committed against that identity, and any regeneration should invalidate approval instead of silently refreshing the artifact. I also think the verifier should fail if the displayed tuple and manifest do not match, so the rule is enforced rather than only documented.