DEV Community

Finley Zhou
Finley Zhou

Posted on

The Evidence Path Is Read-Only for Agent Patches

An agent patch is reviewable only when its evidence path is unchanged. Fixture locks, property oracles, and flake-freeze records are how a reviewer decides that a failure is real. If the same diff may edit those files, the generator can delete the failure instead of fixing the product.

That is the gate. Product code may change. Evidence may not, except in a later human commit that contains no product fix.

Generation is the cheap step. The decision that still needs a procedure is which paths the generator may write, and which results still count after it writes them.

What the gate scores

The failure mode in scope is a green suite produced by editing evidence. The assertion moved, a fixture byte moved, or a freeze line gained a new name. The product bug stayed where it was.

Score three objects. Do not score a single badge.

  1. The fixture lock is a content hash of inputs the test may read.
  2. The property oracle is the check that must still reject a known bad input.
  3. The freeze ledger is the only record that may ignore a flaky test, and only with a citation this patch did not write.

A remote pass is not a fourth object. It is a citation attached to evidence you already stored. A second runner can confirm a fail class. It cannot close a local fixture mismatch, and it cannot mint a freeze inside the agent diff.

Step 1. Classify paths before reading the diff body

Names come first. A summary that says "tests untouched" is not evidence. One field inside tests/fixtures/order.json is enough to invalidate the run.

git fetch origin main
git diff --name-status origin/main...HEAD > /tmp/patch_status.txt
find tests/fixtures -type f -print0 | sort -z | xargs -0 sha256sum > /tmp/fixture.lock
sha256sum -c /tmp/fixture.lock
Enter fullscreen mode Exit fullscreen mode

Read the status column before the patch body. A, M, and D are the rows this sketch parses. A rename row, usually R100, carries two paths and must be classified by hand or by an extended parser.

Map each parsed path to one class. The prefixes are a repo convention, not a universal standard.

Class Rule
product path starts with src/
fixture path starts with tests/fixtures/
oracle path starts with tests/properties/ or tests/oracles/
freeze_ledger path is flake_freeze.json or starts with tests/freezes/
added_test status A under tests/, and not an evidence prefix
other CI YAML, lockfiles, workflows, and anything else

If any path lands in fixture, oracle, or freeze_ledger, stop. Do not spend another model turn explaining the edit. The evidence path is dirty, so the patch is out of scope.

other is a hold, not a pass. A workflow file can raise a timeout and hide a flake without touching product code. Keep those diffs for a human until the repo publishes an allowlist.

Deleting a fixture is the same class of edit as rewriting one. Status D on an evidence prefix is still a reject.

Step 2. Retain the property counterexample

Properties should describe behavior the patch claims to preserve. They should already exist on the base revision. An oracle added in the same diff is not a property. It is an unreviewed claim.

Store a sample even when the check passes. A pass with no retained input is weaker than a result a reviewer can recompute.

def shipping_total(weight_g: int, zone: int) -> int:
    base = 200 + (weight_g // 100) * 40
    return base + {1: 0, 2: 150, 3: 400}[zone]

def zone_monotonic(weight_g: int) -> tuple[bool, dict]:
    totals = [shipping_total(weight_g, z) for z in (1, 2, 3)]
    sample = {"weight_g": weight_g, "totals": totals}
    return totals[0] <= totals[1] <= totals[2], sample
Enter fullscreen mode Exit fullscreen mode

This pair is an unexecuted illustration. It is not a production rate card and not a benchmark. For weight 250, the totals are [280, 430, 680]: base is 200 + (250 // 100) * 40 = 280, then zone offsets add 0, 150, and 400.

Write that sample to tests/properties/samples/zone_monotonic.json and hash the file with the fixture lock. The agent patch may read the sample. It may not rewrite it. If a later function returns a different triple, the retained sample is stale and the lock step should fail. That comparison does not need a model.

Walk only a bounded integer range you are willing to store. Keep the first failing sample if one appears. A failing sample belongs on the evidence path, which means the agent diff still may not edit it in order to go green.

Step 3. Decide from one table

Use the same rows on every patch. The useful part is the reject set, not the happy path.

Paths in the diff Fixture lock Local property Second runner Freeze in this patch Gate
product only match pass, sample retained not required not opened review product diff
any evidence path any any any deny reject
added test match pass any cannot inherit separate review
product only match fail, seed saved same fail class deny in-diff hold; cite later
product only local fail any remote pass deny reject
other (CI, lockfile) match pass any deny hold for a human

Two readings stay fixed because they are easy to optimize away. An added test cannot inherit an old freeze. A remote pass cannot close a local fixture fail or a local property fail. The table keeps those readings beside the path rule so they remain one decision.

A hold is not a soft pass. Nothing merges from a hold. The patch waits until the evidence is clean, or until a human opens a different commit that edits only the evidence path.

Step 4. Encode the contract

The module below is a proposal. It was not executed against a live repository for this article. Review the exit strings before wiring them into CI. Rename rows are intentionally unresolved: the script returns a hold instead of guessing.

import argparse
import sys
from pathlib import Path

PREFIX = {
    "fixture": ("tests/fixtures/",),
    "oracle": ("tests/properties/", "tests/oracles/"),
}

def classify(path: str, status: str) -> str:
    if path == "flake_freeze.json" or path.startswith("tests/freezes/"):
        return "freeze_ledger"
    for label, prefixes in PREFIX.items():
        if any(path.startswith(p) for p in prefixes):
            return label
    if status == "A" and path.startswith("tests/"):
        return "added_test"
    if path.startswith("src/"):
        return "product"
    return "other"

def gate(rows, fixture_match: bool, local_property: str, remote: str) -> str:
    if not rows:
        return "hold: empty diff"
    classes = {classify(path, status) for status, path in rows}
    if classes & {"fixture", "oracle", "freeze_ledger"}:
        return "reject: evidence path edited"
    if "other" in classes:
        return "hold: unclassified path"
    if not fixture_match:
        return "reject: local fixture lock failed"
    if local_property == "fail" and remote == "pass":
        return "reject: remote pass cannot clear a local property fail"
    if "added_test" in classes:
        return "hold: added tests need a separate review"
    if local_property == "fail" and remote == "same_class":
        return "hold: cite the pre-existing test later, do not freeze it here"
    if local_property == "pass" and classes <= {"product"}:
        return "review: product diff only"
    return "hold: incomplete evidence"

def main() -> int:
    parser = argparse.ArgumentParser()
    parser.add_argument("--status", required=True)
    parser.add_argument("--fixture-match", choices=("yes", "no"), required=True)
    parser.add_argument("--local-property", choices=("pass", "fail"), required=True)
    parser.add_argument("--remote", choices=("not_run", "pass", "same_class"), required=True)
    args = parser.parse_args()
    rows = []
    for line in Path(args.status).read_text().splitlines():
        if not line.strip():
            continue
        parts = line.split("\t")
        if len(parts) != 2 or parts[0] not in {"A", "M", "D"}:
            print("hold: rename or unparsed status; classify both paths")
            return 2
        rows.append((parts[0], parts[1]))
    decision = gate(
        rows,
        fixture_match=args.fixture_match == "yes",
        local_property=args.local_property,
        remote=args.remote,
    )
    print(decision)
    return 0 if decision.startswith("review:") else 1

if __name__ == "__main__":
    sys.exit(main())
Enter fullscreen mode Exit fullscreen mode

Path class is checked first. A matching hash does not rescue a dirty oracle. A remote pass does not rescue a local property fail. Empty model output and transport errors stay remote=not_run. Do not coerce them into pass.

python evidence_gate.py \
  --status /tmp/patch_status.txt \
  --fixture-match yes \
  --local-property pass \
  --remote not_run
Enter fullscreen mode Exit fullscreen mode

Accept only the string review: product diff only. Every other string stops the merge. Log it beside the base SHA and the head SHA. A later citation should point at that log, not at a chat transcript.

A compact record shape, with placeholders rather than a claimed run, is enough for review:

{
  "base": "<base sha>",
  "head": "<head sha>",
  "classes": ["product"],
  "fixture_lock": "<sha256 of tests/fixtures tree>",
  "sample": "tests/properties/samples/zone_monotonic.json",
  "local_property": "pass",
  "remote": "not_run",
  "freeze_in_diff": false,
  "decision": "review: product diff only"
}
Enter fullscreen mode Exit fullscreen mode

Step 5. Put free model access and a free server on the correct side

Disclosure: This article was prepared as part of MonkeyCode's product outreach.

MonkeyCode is relevant as a project the operator describes as open source, with free model access and a free server option. Those availability claims are operator-supplied. This article does not add a token allotment, a model name, a hardware shape, a duration, or a benchmark. Those terms move. Read the current project documentation when the job starts, and do not copy a number from an older post.

Use free model access to draft the product diff only. State the writable set in the task: src/ may change; tests/fixtures/, tests/properties/, tests/oracles/, and tests/freezes/ may not. After the draft, run Step 1. If the draft touched the evidence path, discard that draft.

Do not ask the model to restore those files inside the same patch. Restoration is git checkout of the base versions, followed by a fresh classification. A second draft that "puts the tests back" is still an evidence edit if the patch records it.

Use the free server as a second runner only after the local property has failed with a saved sample. Replay that sample. A matching fail class can support a citation for a pre-existing test. It still does not authorize a freeze line in the agent diff. If the server passes while the local lock fails, the table already says reject.

Limitations

Prefix rules miss runtime tampering. Product code can still patch an oracle after import. If helpers live under src/, review the import graph, or load oracles from a path the product package cannot write.

Hash locks miss inputs fetched at runtime. Hash the recorded response body and store that hash on the evidence path. Otherwise the lock checks files the test never read.

Two runners can share one bug and emit the same fail class. A matching citation is not a flake rate. Do not use this gate when the release question is how often a test fails. The gate answers whether this diff edited the files that would hide a fail.

The status parser does not understand copy or rename rows. Treat any rename into an evidence prefix as a reject until the parser grows that case. A green exit from the sketch is not meaningful on a diff that contains R or C rows.

Human oracle repairs are a different lane. If the property is wrong, a person edits it in a commit with no product change. Mixing the repair and the fix is the case this workflow rejects.

Skip the approach when the tree has no fixtures, when the pull request exists to change oracles, or when every test file is new. Added tests need their own review. They are not coverage you can launder through an old freeze.

What to keep

Keep the evidence path read-only for agent patches. Classify status and path, lock fixture bytes, retain property samples, and refuse in-diff freezes. A free model can draft the product change. A free server can add a second citation after a local fail.

Neither one may edit the files that decide whether that citation counts. If that split is already your rule, MonkeyCode's free model access and free server option are enough to rehearse it on a scratch branch. Check the live terms first, then leave the evidence path out of the draft.

Top comments (1)

Collapse
 
suppdevbot profile image
DEV SUPPORTS •
You need to verify your account.
Enter fullscreen mode Exit fullscreen mode

tr.ee/dev-to