DEV Community

Finley Zhou
Finley Zhou

Posted on

Query the Freeze Store Only After a Bound Fixture Miss

An agent patch should not reach a flake-freeze lookup until two cheaper records exist. Those records are a property check that can fail, and a fixture binding that names the generator inputs. Without them, a freeze hit is not evidence of instability. It is evidence that the pipeline asked the wrong store first.

That ordering is the strategy. Property checks reject or clear. Fixture replay then spends compute on a bound input set. The freeze store is a read-only note, consulted last, and its return value is outside the pass set of this runner.

The rest of this piece is a workflow and a small stage machine. The listings are a proposal. They have not been executed against a production fleet, and they do not report measured flake rates.

What each artifact may decide

Three artifacts show up in agent-patch review. They answer different questions, so they should not share one score.

A property check is a predicate over the diff and a small input domain. It may reject, or it may clear. It may not report a flake.

If the predicate stays true for every patch, including an empty diff, it is not a check. Refuse it before any replay process starts. A clear result is only meaningful after that refusal path has been shown to work.

A fixture is a generator plus pinned inputs, not a saved stdout file. The binding fields are patch digest, generator id, inputs digest, and seed. Replaying that tuple can confirm an invariant or record a miss. A matching stdout digest is not the binding, and it is not a certificate.

A flake freeze is a scheduling row: a test target, an owner, and an expiry. It can explain why a replay miss should wait for a person. It should not be queried while the property lane is still empty. A dashboard that left-joins the freeze table onto every run will count missing specs as instability, which sends the on-call to the wrong queue.

Put a budget on each stage

Budgets keep the order from collapsing under time pressure. The figures below are repo policy for a laptop-sized check, not measured optima. Set them in config. Do not treat them as folklore.

Stage Runs when Policy budget Allowed codes Must not
Property Always, first 12 predicates, 2s each CLEAR, REJECT, NOT_FALSIFIABLE Open the freeze client
Fixture Property code is CLEAR 1 bound replay, 30s REPLAY_MATCH, REPLAY_MISS, UNBOUND Hash stdout and call it the fixture id
Freeze lookup Fixture code is REPLAY_MISS 1 read, zero writes ANNOTATED_HOLD, NO_RECORD Return a pass status

If property evaluation returns REJECT or NOT_FALSIFIABLE, exit before the fixture process is spawned. If the fixture is UNBOUND or REPLAY_MATCH, do not construct the freeze client. A miss is the only code that makes a freeze read eligible.

1. Register a falsifier slot

Each property predicate ships with a counterexample probe. The probe is a patch the predicate is required to reject. If the probe is accepted, the predicate is not falsifiable, and the candidate patch is not evaluated.

python -m gate.property register \
  --spec specs/no_bare_except.py \
  --probe probes/bare_except.patch \
  --out .gate/property-slot.json
Enter fullscreen mode Exit fullscreen mode

The slot file stores the spec digest, the probe digest, and the probe result. A slot whose probe did not fail is invalid. Stop there. These commands are local proposals. They assume a gate package you own, not a public installer.

2. Evaluate the candidate against that slot

Run the same spec on the agent patch. Keep the domain narrow: AST predicates, diff bounds, forbidden call sets, schema shape. Persist the reason code on a clear result too, so a later report cannot invent an instability story.

python -m gate.property eval \
  --spec specs/no_bare_except.py \
  --patch candidate.patch \
  --slot .gate/property-slot.json \
  --out .gate/property-result.json
Enter fullscreen mode Exit fullscreen mode

Use three codes, and stop on anything other than PROPERTY_CLEAR. Fixture spend is not a consolation prize for a vague spec.

  1. PROPERTY_REJECT carries a witness path inside the patch.
  2. PROPERTY_CLEAR carries the spec digest and the slot digest.
  3. PROPERTY_NOT_FALSIFIABLE fires when the slot is missing or the probe did not fail.

3. Bind generator inputs, then replay once

A fixture directory without a generator manifest is unbound. The manifest names the generator, the input files, and the seed. Hash those inputs. Do not hash the program stdout and store that hash as the fixture identity.

python -m gate.fixture bind \
  --patch candidate.patch \
  --manifest fixtures/parser/manifest.json \
  --out .gate/fixture-binding.json

python -m gate.fixture replay \
  --binding .gate/fixture-binding.json \
  --invariant invariants/parser_roundtrip.py \
  --out .gate/fixture-result.json
Enter fullscreen mode Exit fullscreen mode

REPLAY_MATCH means the invariant held on that binding. It does not mean the patch is correct off that binding. Keep the binding digest in the result file either way.

REPLAY_MISS stores that same digest next to the failure, so a later run can tell a new miss from a repeated one. UNBOUND means the manifest omitted an input or a seed. That is a pipeline defect. Do not file it as a flaky test.

4. Read the freeze store only after a miss

The lookup key is the test target plus the invariant id. It is not the patch digest. A freeze describes the target history. Keying it to one patch would attach that history to a single edit.

python -m gate.freeze lookup \
  --target parser.roundtrip \
  --invariant invariants/parser_roundtrip.py \
  --require-fixture .gate/fixture-result.json \
  --out .gate/freeze-note.json
Enter fullscreen mode Exit fullscreen mode

The command refuses to start if the fixture result is missing or its code is not REPLAY_MISS. A store hit writes ANNOTATED_HOLD and the freeze id. A store miss writes NO_RECORD, and the runner still holds. Build the client inside this branch only.

5. Publish a stage trace, not a single badge

Reviewers need to see which stage ran. A log line that only says a freeze was applied is a failed pipeline, even when the job is green. Emit one JSON document per run, and upload the .gate directory with the artifact action your forge currently pins. Do not copy an action major version from an old workflow without checking that pin.

{
  "patch_digest": "sha256:example-not-a-real-run",
  "freeze_consulted": true,
  "verdict": "HOLD",
  "stages": [
    {"name": "property", "code": "PROPERTY_CLEAR"},
    {"name": "fixture", "code": "REPLAY_MISS"},
    {"name": "freeze_lookup", "code": "NO_RECORD"}
  ]
}
Enter fullscreen mode Exit fullscreen mode

That document is an example shape. The digest is a placeholder. Elapsed time is omitted on purpose, because this proposal has not been timed.

The runner that produces the shape is also a proposal. Module names are local. Nothing here is a package to install from this page.

# Proposal. Unexecuted. Exit codes follow the budget table.
from enum import Enum
from pathlib import Path
import json
import subprocess

class Stop(Enum):
    CLEAR = 0
    REJECT = 1
    HOLD = 2

def _run(cmd: list[str], out: Path) -> dict:
    subprocess.run([*cmd, "--out", str(out)], check=True)
    return json.loads(out.read_text())

def gate(patch: Path, work: Path) -> tuple[Stop, str]:
    prop = _run(
        ["python", "-m", "gate.property", "eval",
         "--spec", "specs/no_bare_except.py",
         "--slot", str(work / "property-slot.json"),
         "--patch", str(patch)],
        work / "property-result.json",
    )
    if prop["code"] != "PROPERTY_CLEAR":
        return Stop.REJECT, prop["code"]

    bound = _run(
        ["python", "-m", "gate.fixture", "bind",
         "--patch", str(patch),
         "--manifest", "fixtures/parser/manifest.json"],
        work / "fixture-binding.json",
    )
    if bound["code"] != "BOUND":
        return Stop.HOLD, "FIXTURE_UNBOUND"

    replay = _run(
        ["python", "-m", "gate.fixture", "replay",
         "--binding", str(work / "fixture-binding.json"),
         "--invariant", "invariants/parser_roundtrip.py"],
        work / "fixture-result.json",
    )
    if replay["code"] == "REPLAY_MATCH":
        return Stop.CLEAR, "REPLAY_MATCH"
    if replay["code"] != "REPLAY_MISS":
        return Stop.HOLD, replay["code"]

    note = _run(
        ["python", "-m", "gate.freeze", "lookup",
         "--target", "parser.roundtrip",
         "--require-fixture", str(work / "fixture-result.json")],
        work / "freeze-note.json",
    )
    return Stop.HOLD, note["code"]
Enter fullscreen mode Exit fullscreen mode

Call that function from one CI step. Archive the JSON files on both success and failure. If a freeze client was constructed on a PROPERTY_REJECT run, the ordering bug is in the runner, and the verdict is unusable.

6. Debug a HOLD without reopening earlier stages

When the verdict is HOLD, read the last code only. Do not rerun earlier stages unless that code says their artifact is missing.

  1. FIXTURE_UNBOUND: print the manifest and confirm every input path and the seed field. Fix the manifest. Do not open the freeze store to explain the hold.
  2. PROPERTY_BUDGET: the spec exceeded the 2s cap or the 12-predicate cap. Split the spec. Do not raise the cap inside the job to force a green run.
  3. NO_RECORD: the miss is not a known freeze. Keep the hold and assign a person. Do not write a freeze row from the CI job.
  4. ANNOTATED_HOLD: confirm the row expiry is still in the future and the target id matches. If the expiry field is empty, treat the row as absent and file a store defect.

This diagnostic is a reading order for the trace files. It is not a measured incident review.

Draft the predicate off to the side

Writing the first predicate is the slow step. A model can propose a draft from the diff. The falsifier slot still has to reject the probe, so an always-true draft dies in stage 1.

Disclosure: This article was prepared as part of MonkeyCode's product outreach.

MonkeyCode's free model access is a way to draft that predicate. The free server option is a way to execute the draft once before a reviewed file is copied into specs/. Those two availability claims are the only product facts used here. This article does not state quotas, model names, hardware, or how long either option remains available.

An execution on that server is not a fixture binding. The stage machine stays in the repo. A team that already has a reviewed spec library can skip the model step.

After a draft, run two local checks before registration. The probe patch must fail the spec. The empty patch must not fail it, unless the spec is intentionally global. Then write the slot file, and only then point the gate at the agent patch.

Limitations

Order does not repair a wrong invariant. A check can fail the probe and still encode the wrong rule. A bound replay can still miss the input that would have caught the edit.

The time caps also hide slow checks. A 2s property timeout should be a hold with its own code, PROPERTY_BUDGET, or the dashboard will mix unfinished specs with product failures. Raising the cap without splitting the spec only moves the queue.

The lookup is weak on purpose. Some other process must write freeze rows, with an owner and an expiry. This workflow does not define that writer. If the store has no expiry, do not add the lookup. A row that never ages out becomes an allowlist the moment a later change treats ANNOTATED_HOLD as a skip.

The sample does not authenticate the store, lock the manifest against concurrent edits, or shrink a witness. Those controls belong in other changes. Their absence here is not permission to skip them.

Who should skip it

Skip the pipeline when the generator calls live services you cannot pin. An unbound network read fails closed as UNBOUND or flaps as REPLAY_MISS. Neither code describes the patch.

Skip it when the same diff edits the gate, the spec language, or the probe set. The falsifier is then graded by the code under test. Land that change alone, as a toolchain patch.

Skip it when the decision you need is a production rollback. This gate holds or clears a diff. It does not page, revert, or estimate user impact.

Keep three fields in the log

Publish the reason code, the stage that emitted it, and whether the freeze client was constructed. If those three fields are present, a reviewer can audit the budget without replaying the job. If they are absent, the green badge is not an admission record.

Fit the reason-code table to the CI you already operate. If the gap is a place to draft predicates before that local gate runs, MonkeyCode's free model access and free server option are one place to try the draft step. Admission still happens on the local gate.

Top comments (0)