DEV Community

Casey Chen
Casey Chen

Posted on

Model Snapshots Fail as Oracles: A Three-Bucket Review for Agent Diffs

An agent pull request that deletes a local fixture and commits a model transcript has not added a test. It has stored one observation. Keep the intent to cover the parser. Revert the live call, the golden file, and any silent reroute. Prove the behavior with a fixture the repository owns.

That rule does not depend on which assistant drafted the patch. Review the diff. The generator's confidence is not evidence.

What landed in the branch

The shape is usually small. A file such as tests/golden/answer.json appears. A helper builds a client from a base URL. The assertion becomes equality against the new file. CI changes just enough for the job to reach an external inference host.

Sometimes a second hunk hides in the client. On timeout, quota, or HTTP 429, the code assigns a different base URL and sends the same payload again. That is not error handling. It is a second deployment, selected at runtime, with no review of which data leaves the process.

Read network edges before you read the PR prose. A required check that opens a socket will fail for reasons that are not regressions in your code. It also makes every fork run a client of an external host.

Trust, revert, prove

Split every hunk into three buckets before you write a comment. Short diffs still need the split. Length is a weak signal.

Trust lines that remain true when the network is off:

  • A pure parser for the completion payload.
  • The short list of fields the application actually reads.
  • A regression expressed as bytes in the repo, written by a reviewer or generated from a schema you control.

Revert lines that bind the merge gate, or production, to an environment you do not control:

  • HTTP clients in the default test path.
  • Checked-in transcripts used as equality oracles.
  • An except block that swaps in another base URL.
  • Logs of raw prompts, tool outputs, or Authorization values.
  • Hard-coded quotas, model identifiers, regions, or comments that say the host is permanently free.
  • Shell helpers that build a command string and pass shell=True.

Prove with two jobs. The contract job uses a local cassette and is required. The smoke job hits an external host only when someone dispatches it. Smoke may check a status code. It must not check exact prose.

Why equality against model text fails

Three properties of the design break the oracle. No benchmark is required to see them.

Sampling and server-side changes make byte-identical output unstable. A green run today does not pin the server you will get next week. The transcript can also echo the prompt. If the agent copied a credential or a customer string into the request, the golden file publishes it.

So the file is evidence of a past call, closer to a log than to a specification. Logs are not contracts. Merge the parser if it is sound. Delete the captured answer.

Latency fields fail the same way. A cassette that asserts elapsed_ms < 800 couples your build to someone else's queue. Drop timing, token-usage, and provider metadata from any file you intend to keep.

Cassette rules

A cassette is a contract only if a reviewer can explain every field. Keep three ideas in the file: a synthetic input label, the payload fragment the parser accepts, and the expected structure. Drop chain-of-thought text and raw tool traces. Those values change without your application code changing, and some of them are sensitive even when the prompt looked harmless.

Name the file after the behavior, not after the host or the date. parse_completion.json survives a server change. free_host_2026_10_10.json does not. If the agent used a timestamped name, rename it in the same review. The filename is part of the claim.

Worksheet

Fill the action from the patch, not from the description.

Question Fail signal Action
Does required CI open a socket? httpx, requests, curl, or a client import under tests/ without a skip mark Move it out of the default job
Is model prose committed? New JSON or text under tests/ or fixtures/ whose body is a completion Revert; add a hand-written cassette
Can cache keys collide across tenants? Key derived only from the prompt text Revert, or include tenant id and contract version
Does an error change the host? except assigns a second base URL Revert silent failover; default the flag off
Did secrets enter the diff? .env, Bearer, sk-, full prompt dumps Revert; rotate if the commit is already public
Does the PR state a quota or model name it cannot cite? Token counts, hardware, permanence Delete the claim until a primary source exists

An external inference option is not an allowlist, a capacity plan, or an oracle. If the patch treats it as any of those, that hunk is a revert.

Commands

These commands are a proposed review sequence. They are not output captured from a named repository.

git diff --stat origin/main...HEAD
git diff origin/main...HEAD -- tests .github docker-compose.yml
rg -n -i "httpx|requests|curl|base_url|subprocess|shell=True|golden" tests .github
rg -n -i "authorization|api_key|bearer |sk-|prompt" -g "!**/node_modules/**"
Enter fullscreen mode Exit fullscreen mode

Open new files under tests/ first. The risky call is often in a helper with a harmless name, such as client_for_tests or load_golden.

If the search hits shell=True, stop on that hunk. A shell command is concatenation. Unquoted expansion splits on spaces and interprets metacharacters. Prefer a fixed argument list and shell=False, or revert the helper. A note that the agent already ran the script is not a review outcome.

Proposed harness

The Python below is a sketch for reviewers to demand. It has not been executed for this article. It is not a benchmark, and the helper probe_status is a placeholder.

import json
import os
from pathlib import Path

import pytest

CASSETTE = Path(__file__).parent / "cassettes" / "parse_completion.json"

def parse_completion(payload: dict) -> dict:
    choice = payload["choices"][0]["message"]["content"]
    return {"text": choice, "truncated": False}

def test_parser_reads_contract_fields_only():
    payload = json.loads(CASSETTE.read_text(encoding="utf-8"))
    assert parse_completion(payload)["text"] == "cassette-text"

@pytest.mark.smoke
def test_optional_reachability_smoke():
    pytest.importorskip("httpx")
    base = os.environ.get("SMOKE_BASE_URL")
    if not base:
        pytest.skip("SMOKE_BASE_URL unset; smoke is opt-in")
    status = probe_status(base)  # proposal: GET {base}/models or /health
    assert status == 200
Enter fullscreen mode Exit fullscreen mode

The cassette should contain boring synthetic text. Do not paste a live completion into it. Exact prose assertions belong only on that local file, because you can change the file in the same pull request as the parser.

Pair the tests with a workflow fragment. This fragment is also unexecuted.

jobs:
  contract:
    steps:
      - run: pytest -m "not smoke"
  smoke:
    if: github.event_name == 'workflow_dispatch'
    steps:
      - run: pytest -m smoke
        env:
          SMOKE_BASE_URL: ${{ vars.SMOKE_BASE_URL }}
Enter fullscreen mode Exit fullscreen mode

Manual dispatch is the point. An external host on the automatic path becomes a flaky required check. It is also an unbounded client of shared capacity whenever the workflow retries.

What not to multiply

Retries, matrix builds, and parallel smoke jobs multiply calls. A retry policy that was reasonable for a paid primary can be a load spike against a shared server. Keep smoke to one probe and no retry loop.

If you need throughput numbers, that is a different test. It needs an owner, a concurrency cap, and a stop time. This worksheet does not define that test, and a green status code does not supply those numbers.

Cache refills have the same shape as a bad golden file. A job that warms a cache by replaying production prompts at an external host is a data export. Refuse it unless every prompt is synthetic and the cache key includes a tenant boundary plus a contract version.

Using a free model path without promoting it

Disclosure: This article was prepared as part of MonkeyCode's product outreach.

MonkeyCode enters this review in two narrow places. Free model access can help a reviewer draft the comment or a synthetic cassette, after which a human still reads the diff. A free server option can be the value of SMOKE_BASE_URL on the manual job.

That is the full scope used here. This draft does not name models, token quotas, hardware, duration, or permanence, because those details were not supplied as citable sources. If a generated README claims a specific allowance, do not publish the figure. Ask for a primary document first. Availability is not a capacity plan.

Delete the product name and the worksheet still holds. Observations are not contracts. Required checks should not depend on an external inference host.

Who should skip this

Skip the split if you cannot keep fixtures free of customer data. Synthetic prompts only. If a real prompt is already in the branch, reverting that file is the first action. Rotate any credential that sits beside it.

Skip it if CI cannot exclude marks. A smoke test that still runs on every pull request is the original bug with a new name. Fix the workflow filter before discussing hosts.

Skip it when the main change is authorization, tenancy, or a tool that has side effects. Those patches need a threat model. A cassette will not catch a confused-deputy bug in a shell tool, and a status code will not show which files the tool touched.

Streaming responses also sit outside the sketch. Partial buffers, disconnects, and reassembled tokens need their own assertions. Do not freeze one stream mid-flight and call that file the spec.

Comment shape

Write the three buckets in the review. A usable comment looks like this:

  1. Trust the parser and the field list in the application module.
  2. Revert the golden transcript, the default base URL in tests, and the failover assignment.
  3. Prove behavior with the cassette on the required job. Leave external reachability on manual dispatch.

If a free model path and a free server are already available to the team, attach them only to that manual job, after the required check is local. The merge decision stays on the cassette either way.

Top comments (0)