DEV Community

Finley Zhou
Finley Zhou

Posted on

Changed Symbols Earn Property Credit, or the Agent Patch Stays Closed

An agent patch does not merge because the job is green. It merges when every changed symbol carries property credit, every fixture hash matches the lock, and every flake freeze contributes zero.

CI collapses three unlike signals into one bit. A property can pass. A fixture can replay. A known flake can be silenced. Those outcomes are not substitutes. The miss this ledger targets is specific: production behavior changed, no invariant covers the edited symbol, and a freeze is what kept the check quiet.

The scoring rules below are a proposal. The classifier is an unexecuted example, not a measured gate from a production repository.

What receives credit

Property credit is the only positive unit. A changed symbol earns one point when a predicate states an invariant, the predicate's source hash is stored in the receipt, and the predicate does not compare the result to a frozen document.

Fixture locks are mandatory and worth nothing. A matching hash shows that the input bytes are the bytes you reviewed. It does not show that the behavior is acceptable.

A flake freeze is a quarantine citation, also worth nothing. It may point at a baseline flake so an unrelated failure does not block review. It may not stand in for a missing predicate on a symbol this patch edited.

Coverage is credited symbols divided by changed symbols. Admission is exact, not approximate.

admit  if coverage == 1.0
       and fixture_status == "match"
       and rejection_codes == []
reject otherwise
Enter fullscreen mode Exit fullscreen mode

A coverage of 0.5 is not a soft warning. The receipt rejects, and the missing symbols go back to the author.

Where a Boolean gate stays silent

Agent diffs often include a new test next to the fix. That pattern looks careful. It fails the ledger when the assertion equals a fixture the same diff rewrote, or when a skip cites a freeze id that the patch itself created.

The job still passes. The receipt does not. R5 marks fixture equality dressed up as a check. R4 marks a freeze id absent from the baseline ledger. R3 marks a baseline freeze used as the only signal on a changed symbol.

A quieter miss is the helper that returns true after reading fixture state. The test name contains the word property. The predicate hash moves whenever the fixture moves, so the classifier withholds credit. A label in the test name is not an invariant.

Decision table

Lane Credit Admit rule Remote run allowed as proof? Rejection
Property on a changed symbol 1 per symbol Coverage must be 1.0 Only as confirmation after a local pass R1 missing property, R5 fixture equality
Fixture content hash 0 Exact match required No R2 hash mismatch
Baseline flake freeze 0 Not required No R3 freeze as oracle, R4 unknown freeze id
Remote-only log 0 Not accepted No R6 remote-only receipt

Read the table as the policy. The steps below only compute it.

Step 1: List changed symbols from the diff

Start from the patch, not from whichever tests the agent happened to touch. Count a symbol when its definition changes, or when a call site change alters an observable result.

git diff --unified=0 origin/main...HEAD -- '*.py' \
  | awk '/^diff --git/{f=$0} /^@@/{print f, $0}'
Enter fullscreen mode Exit fullscreen mode

Write one qualified name per line into changed_symbols.txt. An empty list ends the review early. A test-only edit can still carry a property if it claims to lock behavior, but it does not earn fixture credit for production files the diff never touched.

Keep the list short enough to review by hand. If a generated file explodes the symbol count, exclude that path in the command and say so in the receipt. Hidden exclusions are just another Boolean.

Step 2: Classify each touched assertion

Label every assertion in a touched test with one class.

  1. property — a predicate over inputs and outputs that does not embed a full expected document.
  2. fixture_eq — equality or snapshot comparison against a checked-in blob.
  3. freeze_ref — skip, xfail, or an annotation that cites a flake id.
  4. opaque — helpers, mocks, or assertions whose meaning is not visible in the diff.

Only class 1 can add credit. Class 2 may satisfy the fixture gate when the blob hash matches. Class 3 never adds credit. Class 4 adds nothing, and any changed symbol that has only class 4 evidence remains uncovered.

The vocabulary holds, implies, and invariant in the sample classifier is a local convention, not a language feature. If your suite uses different predicate names, change the hint before you trust a clearance.

Step 3: Hash fixture bytes, not paths

Path strings survive edits that replace the file. An agent can rewrite tests/fixtures/invoice.json and leave the path in the test unchanged.

python - << 'PY'
import hashlib, pathlib
root = pathlib.Path("tests/fixtures")
for path in sorted(root.rglob("*")):
    if path.is_file():
        digest = hashlib.sha256(path.read_bytes()).hexdigest()
        print(f"{digest}  {path.as_posix()}")
PY
Enter fullscreen mode Exit fullscreen mode

Diff that listing against fixture.lock on the target branch. Any mismatch is R2. Updating the lock inside the patch is allowed only when a property still covers every changed symbol without using the new blob as the expected value.

If the fixture directory is absent, record fixture_status as not_applicable rather than match. Do not invent a hash to satisfy the gate.

Step 4: Cite freezes only from the baseline ledger

The quarantine file must already exist on the target branch. A patch-local freeze file is not a source of valid ids.

flakes:
  - id: FLK-1044
    signature: "TimeoutError: checkout_worker"
    opened_on: "2026-09-02"
    expires_on: "2026-10-16"
    tests: ["tests/test_checkout.py::test_worker_drain"]
Enter fullscreen mode Exit fullscreen mode

Those dates are example fields for the rule, not a report of a live incident. A citation is valid only when the id is in that baseline file, the signature still matches, and expires_on is later than the review clock. On a review clock of 2026-10-11, the example id above is still citable. An id that expired on 2026-10-01 is not.

Do not extend expiry in the behavior patch. Expiry is a policy edit and belongs in its own review. A freeze that merely keeps a changed symbol quiet is R3, even when the id is old and unexpired.

Step 5: Worked receipt

Take two changed symbols, billing.prorate and billing.invoice_total. The first test is assert holds(prorate(10, 0.5) >= 0). That is one credit. The second test is assert invoice_total(order) == load_fixture("invoice.json"). That is R5 and zero credit. An unrelated citation of baseline FLK-1044 adds no credit either.

{
  "changed_symbols": ["billing.prorate", "billing.invoice_total"],
  "credits": {"billing.prorate": 1, "billing.invoice_total": 0},
  "coverage": 0.5,
  "fixture_status": "match",
  "freeze_citations": ["FLK-1044"],
  "rejection_codes": ["R1", "R5"],
  "admit": false
}
Enter fullscreen mode Exit fullscreen mode

The repair is a predicate that does not read the fixture as an oracle, for example assert holds(invoice_total(order) == sum(line.net for line in order.lines)). After that predicate is local-green, coverage can reach 1.0 and R1 clears. R5 stays if the fixture equality remains in the file. Delete it, or keep it only as a non-oracle lock outside the credit column.

A rerun that stays green without the new predicate does not change this receipt. Green is not coverage.

Step 6: Classifier you can adapt

This module is unexecuted proposal code. It will mis-fire on suites that do not use the hint vocabulary. Run it on a scratch branch and compare its codes with a manual read before you wire it to merge.

#!/usr/bin/env python3
"""Unexecuted proposal: symbol credit for an agent-patch receipt."""

from __future__ import annotations

import hashlib
import json
import re

PROPERTY_HINT = re.compile(r"^\s*assert\s+(implies|invariant|holds)\b", re.M)
EQ_HINT = re.compile(r"assert\s+.+\s*==\s*load_fixture\b")
FREEZE_HINT = re.compile(r"flake-id:\s*(FLK-\d+)")

def sha256_text(text: str) -> str:
    return hashlib.sha256(text.encode("utf-8")).hexdigest()

def classify_test(source: str) -> str:
    if EQ_HINT.search(source):
        return "fixture_eq"
    if PROPERTY_HINT.search(source):
        return "property"
    if FREEZE_HINT.search(source):
        return "freeze_ref"
    return "opaque"

def score(changed: list[str], tests: dict[str, str], baseline_flakes: set[str]) -> dict:
    credits = {name: 0 for name in changed}
    codes: list[str] = []
    citations: list[str] = []
    for symbol, source in tests.items():
        kind = classify_test(source)
        found = FREEZE_HINT.findall(source)
        citations.extend(found)
        for flake_id in found:
            if flake_id not in baseline_flakes:
                codes.append("R4")
        if symbol not in credits:
            continue
        if kind == "property":
            credits[symbol] = 1
        elif kind == "fixture_eq":
            codes.append("R5")
        elif kind == "freeze_ref" and credits[symbol] == 0:
            codes.append("R3")
    if any(credit == 0 for credit in credits.values()):
        codes.append("R1")
    coverage = 0.0 if not changed else sum(credits.values()) / len(changed)
    reject = sorted(set(codes))
    return {
        "changed_symbols": changed,
        "credits": credits,
        "coverage": coverage,
        "predicate_hashes": {name: sha256_text(body) for name, body in tests.items()},
        "freeze_citations": sorted(set(citations)),
        "rejection_codes": reject,
        "admit": coverage == 1.0 and not reject,
    }

if __name__ == "__main__":
    sample = score(
        changed=["billing.prorate", "billing.invoice_total"],
        tests={
            "billing.prorate": "def test_prorate():\n    assert holds(prorate(10, 0.5) >= 0)\n",
            "billing.invoice_total": "def test_total():\n    assert invoice_total(order) == load_fixture('invoice.json')\n",
        },
        baseline_flakes={"FLK-1044"},
    )
    print(json.dumps(sample, indent=2))
Enter fullscreen mode Exit fullscreen mode

The sample input is the worked case from Step 5, encoded as strings so the file runs without a repository. Replace those strings with real test bodies before you treat the output as a review artifact.

python tools/merge_ledger.py > receipt.json
python -c 'import json,sys; r=json.load(open("receipt.json")); sys.exit(0 if r["admit"] else 1)'
Enter fullscreen mode Exit fullscreen mode

The process exit status is the mechanical gate. A reviewer still reads rejection_codes before any override. Overrides belong in the receipt as an explicit field, and they cannot convert freeze citations into credit.

The scorer above does not compute fixture_status, and it does not parse expiry. Intersect its admit flag with the Step 3 lock comparison, and feed it only flake ids that Step 4 already accepted. A true flag with a hash mismatch is still R2.

Step 7: Draft a missing predicate, then prove it locally

Disclosure: This article was prepared as part of MonkeyCode's product outreach. MonkeyCode's free model access has one job here. After an R1 receipt lists uncovered symbols, you can ask a free model to draft candidate assert holds(...) lines from the diff. Those lines are untrusted text.

Keep a draft only if you can restate the invariant without the draft in front of you, and only after a local run executes it on the patched code. A draft that merely restates the fixture file is class fixture_eq, not credit.

The free server option has a second, separate job. Once the local property lane is green, rerun that pure lane remotely and attach the log as confirmation. Confirmation is not credit. It does not clear R2, R3, R4, or R6.

A receipt that contains only a remote log is rejected. Both options are operator-supplied availability claims for this workflow. This article does not state a model name, quota, hardware profile, duration, or permanence. If either option is unavailable, compute the receipt locally. admit must not flip because a remote lane was missing.

Step 8: Read four fields, in order

  1. coverage equals 1.0. Anything lower returns to the author with the uncovered names.
  2. rejection_codes is empty. A lingering R5 means a fixture equality is still in the diff, even if some other predicate scored a point.
  3. Each id in freeze_citations is baseline, signature-matched, and unexpired against the CI clock printed in the receipt.
  4. predicate_hashes match the test bodies in the diff. A hash that moved without a reviewed edit is a stale receipt, and stale receipts do not admit.

Merge only after those four checks. Do not lower the coverage bar to "most symbols" to clear a queue.

Limitations

The classifier is lexical. Helper-shaped properties become opaque and produce false R1 results. Vacuous forms such as assert holds(True) can look like credit unless you add a constant check. Add that check before clearance means anything.

Snapshot-heavy suites will collect R5 on almost every agent patch. That is a prompt to extract one invariant per change, not a prompt to disable the code. If you cannot name an invariant, this ledger does not fit that change.

Expiry uses the clock you pass in. A drifted clock admits dead freezes. Pin the clock to the CI timestamp and print it beside freeze_citations.

Nothing here is a formal proof. A fluent predicate can still encode the wrong business rule. Load, migration, and security reviews stay outside the receipt.

Who should skip it

Skip the ledger on a mechanical rename that your policy already treats as non-behavioral. Also skip it when the edited code has no deterministic predicate without live third-party state. In that setting the strategy manufactures R1 noise, and a recorded interaction contract is the honest artifact instead.

Do not adopt the ledger if you need the freeze file to grant merge credit. That need contradicts the rule. Do not use a free-model draft as the source of record, and do not file a free-server log as the only pull-request evidence.

Regulated teams should keep the local receipt, the lock file, and the CI timestamp in the archive their policy already names. On the next agent patch, start from the rejection codes rather than from a runner brand. The codes are the strategy. Remote confirmation stays optional.

Top comments (0)