DEV Community

Finley Zhou
Finley Zhou

Posted on

Domain, Assertions, Fixture Hash: Gate Agent Patches on Strength

An agent patch that shrinks a property domain, drops an assertion, swaps out a case, or replaces a fixture hash has weakened the test. That delta is a review failure. It is not intermittent, and it is not eligible for a flake freeze.

This note is a pre-merge ledger for that cut. Compare a base property card with the card after the patch, then refuse the change when strength moves the wrong way. Freeze records stay downstream. They never repair a smaller check.

What the ledger measures

A property check fails when an input inside a declared domain violates an invariant. Strength here is not a subjective score. It is a handful of fields recomputed from source and fixture bytes.

Domain size is either the length of a concrete case list or the product of inclusive integer bounds written as literals. Assertion count is the number of assert statements inside the property function, read with ast, not from a test process. Fixture identity is the SHA-256 of the fixture file that property loads. Case identity is the set of case keys, not merely the length of the list.

A patch may add cases or add assertions. It may not remove a base case, narrow a literal bound, or drop an assert unless a named waiver file is present. A fixture hash change is its own class. It does not ride along with a production-code edit.

CPython started with -O does not execute assert statements. Counting them from an optimized run under-reports strength. The ledger reads the file a reviewer can open, and it walks only the property function so helper asserts in the same module do not inflate the count.

Step 1. Snapshot base cards from the parent revision

Check out the parent of the agent patch. Emit one JSON card per property id. Store the cards with the CI artifacts for that SHA. Do not hand-edit them after the snapshot.

git rev-parse HEAD > .ledger/base_sha.txt
python -m strength_ledger snapshot \
  --tests tests/properties \
  --fixtures tests/fixtures \
  --out .ledger/base_cards.json
sha256sum tests/fixtures/balances.json
Enter fullscreen mode Exit fullscreen mode

strength_ledger is a proposed module layout, not a published package. The commands show the interface. They are not a transcript from a live repository, and no pass rate is claimed for them.

Step 2. Snapshot the patch on a throwaway worktree

Apply the agent patch only inside a detached worktree. Re-run the same snapshot. Diff by property id. Formatting noise is irrelevant. A missing id is a deletion, which removes the whole property and fails the gate.

git worktree add --detach ../patch-review "$PATCH_SHA"
python -m strength_ledger snapshot \
  --tests ../patch-review/tests/properties \
  --fixtures ../patch-review/tests/fixtures \
  --out .ledger/patch_cards.json
python -m strength_ledger diff \
  --base .ledger/base_cards.json \
  --patch .ledger/patch_cards.json \
  --out .ledger/strength_delta.json
echo "gate_exit=$?"
Enter fullscreen mode Exit fullscreen mode

Exit 0 means shared properties did not weaken. Exit 2 means a drop, a case swap, a fixture swap, or a removed id. Exit 3 means shared ids held and the patch added new ids. New ids need a separate review path. This ledger does not define how, or whether, any later freeze file may mention them.

Step 3. Compute the delta in a small module

The module below is an unexecuted example. It documents the decision procedure. It does not report a measured run against an agent, a model, or a CI history.

import ast
import hashlib
import json
from dataclasses import dataclass

@dataclass(frozen=True)
class PropertyCard:
    prop_id: str
    domain_size: int
    assertion_count: int
    fixture_sha256: str
    case_keys: tuple[str, ...]
    bounds: dict[str, tuple[int, int]]

def fixture_hash(blob: bytes) -> str:
    return hashlib.sha256(blob).hexdigest()

def assertion_count(source: str, fn_name: str) -> int:
    tree = ast.parse(source)
    for node in tree.body:
        if isinstance(node, ast.FunctionDef) and node.name == fn_name:
            return sum(isinstance(n, ast.Assert) for n in ast.walk(node))
    raise ValueError('missing property function')

def domain_size(cases: list, bounds: dict[str, tuple[int, int]]) -> int:
    if cases and bounds:
        raise ValueError('use a case list or literal bounds, not both')
    if cases:
        return len(cases)
    if not bounds:
        raise ValueError('no cases and no literal bounds')
    size = 1
    for low, high in bounds.values():
        if high < low:
            raise ValueError('inverted bound')
        size *= high - low + 1
    return size

def cases_weakened(base_keys: tuple[str, ...], patch_keys: tuple[str, ...]) -> bool:
    return not set(base_keys).issubset(set(patch_keys))

def bounds_weakened(base: dict, patch: dict) -> bool:
    if set(base) - set(patch):
        return True
    for key, (low, high) in base.items():
        plow, phigh = patch[key]
        if plow > low or phigh < high:
            return True
    return False

def classify(base: PropertyCard, patch: PropertyCard) -> str:
    if base.prop_id != patch.prop_id:
        raise ValueError('cards must share prop_id')
    dropped = (
        patch.domain_size < base.domain_size
        or patch.assertion_count < base.assertion_count
        or patch.fixture_sha256 != base.fixture_sha256
        or cases_weakened(base.case_keys, patch.case_keys)
        or bounds_weakened(base.bounds, patch.bounds)
    )
    if dropped:
        return 'strength_drop'
    same = (
        patch.domain_size == base.domain_size
        and patch.assertion_count == base.assertion_count
        and patch.case_keys == base.case_keys
        and patch.bounds == base.bounds
    )
    return 'unchanged' if same else 'strength_hold_or_up'

def diff_ids(base: dict[str, PropertyCard], patch: dict[str, PropertyCard]) -> int:
    removed = sorted(set(base) - set(patch))
    if removed:
        print(json.dumps({'decision': 'property_removed', 'ids': removed}))
        return 2
    shared = sorted(set(base) & set(patch))
    bad = [i for i in shared if classify(base[i], patch[i]) == 'strength_drop']
    if bad:
        print(json.dumps({'decision': 'reject', 'ids': bad}))
        return 2
    fresh = sorted(set(patch) - set(base))
    if fresh:
        print(json.dumps({'decision': 'new_ids_separate_review', 'ids': fresh}))
        return 3
    print(json.dumps({'decision': 'strength_ok', 'ids': []}))
    return 0
Enter fullscreen mode Exit fullscreen mode

domain_size accepts a concrete case list or inclusive integer literals, never both. Float bounds, sampled generators, and bounds loaded from config are parse errors. A loud error is safer than a guessed product.

Length alone is not enough. cases_weakened fails the patch when any base case key disappears, even if the agent appends new keys and the length rises. bounds_weakened fails the patch when a named interval is narrowed or removed. Extra keys and wider literal intervals count as hold-or-up only when every base input remains included.

assertion_count ignores pytest.raises and custom check_* helpers. Those predicates still matter. They need a project-specific visitor or a human pass. The standard-library count is the minimum bar, not the whole oracle.

Step 4. Apply the decision table before any freeze discussion

Observed delta Domain Assertions Cases and bounds Fixture hash Gate Freeze file
Production edit only same or up same or up base set retained same continue not yet; strength is only the first gate
Case key removed any any base key missing any reject not allowed
Literal bound narrowed down any interval shrinks any reject not allowed
Assert removed any down any any reject not allowed
Fixture bytes changed any any any different reject until fixture review not allowed
Property id removed n/a n/a n/a n/a reject not allowed
New property id n/a n/a n/a n/a separate review not attached here
Waiver file present down, with spec link down, with spec link may shrink may change human merge still not a freeze

A waiver is a checked-in Markdown note. It names the property id, the old card, the new card, and a spec link. It is not a timeout. It does not decay into a green check. Absent that file, exit 2 stands.

Use this card shape. The zeros and empty collections are stand-ins so the schema is visible, not observed CI data. Replace them with snapshot output.

{
  "prop_id": "prop_balance_non_negative",
  "domain_size": 0,
  "assertion_count": 0,
  "fixture_sha256": "<sha256sum of the fixture file>",
  "case_keys": [],
  "bounds": {}
}
Enter fullscreen mode Exit fullscreen mode

Recompute the digest with sha256sum and compare it to fixture_sha256. If they differ, the card is stale. Fail closed. Do not update a freeze list to paper over a stale hash.

Step 5. Let the free model propose inputs, never expected values

Disclosure: This article was prepared as part of MonkeyCode's product outreach.

The availability claims used here are only those supplied for this draft: free model access, and a free server option. No model id, quota, region, hardware shape, or retention period is assumed. None of those belong in the gate.

Free model access has one job, and only after exit 0. It may propose candidate inputs for a property whose card did not weaken. The local property computes the verdict. Model text is not an expected value, and it is not written back into the fixture. A fixture write would change the hash and return the patch to exit 2.

def accept_candidate(raw: str, parser, prop) -> str:
    try:
        value = parser(raw)
    except ValueError:
        return 'discard_unparseable'
    held = prop(value)  # local invariant; raises if the candidate breaks it
    if held is False:
        raise AssertionError('property returned false')
    return 'candidate_held'
Enter fullscreen mode Exit fullscreen mode

Parse once. If parsing fails, discard the string. Do not loop the model until a flattering input appears. One discarded proposal is a valid ledger outcome. Ask the model nothing about whether the patch is correct. That sentence is not a field on the card.

Step 6. Replay a pinned candidate file after the ledger passes

The free server is a second execution site for the same parsed candidates. It is optional. Start it only after strength_delta.json records strength_ok. A rejected patch never reaches that process, so remote availability cannot change a strength decision.

Pin the candidate list before any replay. Include the parent SHA and the property id in the file. Do not regenerate candidates between runs. Two runs with two different prompts are not the same experiment.

python -m strength_ledger replay \
  --cards .ledger/patch_cards.json \
  --candidates .ledger/candidates.json \
  --runner local

# optional; omit this process if the free server cannot be reached
python -m strength_ledger replay \
  --cards .ledger/patch_cards.json \
  --candidates .ledger/candidates.json \
  --runner remote
Enter fullscreen mode Exit fullscreen mode

A candidate that fails the local property is a counterexample. Save it beside the delta JSON and send the patch back for a code fix. That path does not open a flake freeze. A freeze remains a later control for a failure that will not bind to a saved candidate and that sits on an unchanged strength card. How that later record is cited, expired, or reverted is outside this ledger.

Step 7. Keep four review files, not a new platform

Reviewers should see a short artifact set.

  • .ledger/base_cards.json holds parent cards. CI regenerates them. Humans do not edit them.
  • .ledger/strength_delta.json holds the gate output for the patch SHA.
  • .ledger/candidates.json holds parsed inputs, and only after exit 0.
  • docs/waivers/<prop_id>.md is optional, and required when a card is allowed to shrink on purpose.

If those files disagree, the snapshot is wrong. Re-run it. Do not repair the disagreement by appending the property id to a freeze file.

Why a strength drop cannot be filed as a flake

A flake freeze is an exception for a failure that moves while the property source, the fixture bytes, and the candidate list stay fixed. A strength drop is the opposite. The source or the fixture changed, and the change made the check smaller or replaced a base case. The signature is stable: counts, a case set, bounds, and a hash.

Order stays fixed for that reason. Ledger first. Local property execution second. Remote replay third, and only as a duplicate of the pinned file. Freeze review last, and only if the card is flat. This note stops at the eligibility cut. It does not specify freeze prose, runner pairing, or revert policy.

Limitations, and who should skip this gate

The ledger counts structure, not meaning. An agent can preserve assertion count and still replace assert balance == expected with assert isinstance(balance, int). Pair the gate with a predicate reading. A green delta is not proof that the oracle stayed strong.

Equal length can still hide duplicate cases if case keys are not stable. The subset check works only when keys name inputs, not when every run mints a fresh label. If the suite has no stable case key, fail the snapshot. Do not synthesize keys from output text.

Generative tests whose bounds are computed at runtime will not yield a domain size. Fail the snapshot when a bound is not a literal. Do not impute a size from a previous run. Imputed sizes become fake history the moment the generator changes.

Teams that shrink domains every week without a spec link should not turn this gate on. It will block their normal edits. The predictable workaround is a freeze file that launders the shrink. That workaround is the failure this procedure exists to stop. Fix the waiver habit first, or do not adopt the exit code.

Screenshot diffs, prose snapshots, and one-off scripts have no property inventory to card. The module will not invent one. This gate is also the wrong tool when the goal is to silence a red build. Silencing skips the counts. Suites that cannot keep base_cards.json next to the SHA should stay on ordinary review.

Free-model output can fall outside the declared domain. accept_candidate must discard it. Unbounded regeneration is not part of the method. Free-server downtime is not a blocker: record the local replay and move on. Do not wait for a remote process to bless a diff the ledger can already classify. If quotas or model identity matter to the pipeline, confirm them with the operator before encoding them in CI. This article does not list them.

What to run on the first property

Select one property file and one fixture whose hash you can recompute. Generate the base card. Build a second card that deletes one assert or one base case key. The contract check below is part of the proposal. It holds only if the module implements the exit codes above.

python -m strength_ledger diff \
  --base .ledger/base_cards.json \
  --patch .ledger/weakened_cards.json \
  --out .ledger/strength_delta.json
status=$?
test "$status" -eq 2
Enter fullscreen mode Exit fullscreen mode

Restore the deleted line, diff again, and require exit 0 before any candidate file is written. If that pair of exits is stable, candidate drafting can use the free model access already disclosed, and a second replay can use the free server. Neither one replaces the card.

Top comments (0)