A patch-scoped flake freeze is valid only when the agent patch owns the intermittency. If the reverted tree is mixed on the same witness, the freeze belongs on the harness, not on the patch. A second green run does not settle that ownership question.
Agent patches fail property checks for reasons that look identical in a CI log. A stable logic miss, a dirty local fixture, and a harness race all print one red line. Waiving every red line hides regressions, and waiving none of them blocks merges on noise. Ownership has to be measured, which is what a revert control is for.
Close the witness bundle first
A witness bundle is a closed record. It is not a pasted traceback. The checker later in this article refuses incomplete records before it applies any freeze policy. Treat the listing as a reference implementation you can run locally. This article does not report a corpus pass rate, a flake frequency, or a production incident count.
Require these fields before classification starts:
-
witness_id: hash of the property name, the seed, and the raw input bytes that failed. -
fixture_digest: digest of fixture bytes after setup and before the property runs. Include paths in the hashed listing so renamed files do not collide. -
patch_id: checksum of the agent diff you will replay. -
revert_id: tree id of the parent, or of that diff with the agent hunks removed. -
runs: one boolean per planned repeat, tagged withrunner_idandtree(patchorrevert). -
k: repeats per cell, chosen before the first rerun. With two runners and two trees, the run list length must be4 * k.
Runner A is a locked local worktree. Runner B is a clean checkout on another machine or another container. Shared caches, editor swap files, and leftover services disqualify runner A even when the digest string matches.
Use a clean second runner, not a model vote
Disclosure: This article was prepared as part of MonkeyCode's product outreach.
MonkeyCode's free server option is one way to host runner B when the team has no spare clean machine. Free model access can draft a candidate property sentence from the failing input. That sentence is a draft. It counts only after a person turns it into an executable check and commits the code. Model text never occupies a cell in the matrix below. This article states those two availability facts only. It does not name models, quotas, hardware, durations, or permanence.
Drop the product names and the contract is unchanged. Any clean second runner, plus a human-owned executable property, is enough.
Classify cells, then compare trees
Finish every planned run. Do not stop when a cell turns green.
- Mark a cell
STABLE_PASSwhen allkruns passed. - Mark it
STABLE_FAILwhen allkruns failed. - Mark it
MIXEDwhen at least one run passed and one failed. - Accept a tree class only when both runners assign that same label.
- If the runners disagree, or the fixture digests differ, stop with environment skew. Do not average the booleans into a pass rate.
One pattern allows a patch-scoped freeze. Both runners must mark the patch cell MIXED, and both must mark the revert cell STABLE_PASS. The parent property held. The agent patch is the first tree that became intermittent.
Every other agreement pattern refuses that freeze:
- Revert is
MIXEDon both runners. The flake is already in the parent tree, so the target is the harness or the fixture. A waiver attached topatch_idwould hide that. - Revert is
STABLE_FAIL. The parent already misses this witness. The agent change did not create the mode you are about to waive. - Patch is
STABLE_FAIL. Keep the check closed. This is a miss, not an intermittent label. - Patch is
STABLE_PASS. There is no failure to freeze.
Reference checker
Save this as witness_freeze.py. It is proposed code for review. Running the demo prints decision names for four constructed matrices. Those prints are not measurements.
from __future__ import annotations
from collections import defaultdict
from dataclasses import dataclass
from enum import Enum
class Cell(str, Enum):
STABLE_PASS = "STABLE_PASS"
STABLE_FAIL = "STABLE_FAIL"
MIXED = "MIXED"
EMPTY = "EMPTY"
class Decision(str, Enum):
PATCH_FREEZE_CANDIDATE = "PATCH_FREEZE_CANDIDATE"
HARNESS_NOT_PATCH = "HARNESS_NOT_PATCH"
PARENT_ALREADY_FAILS = "PARENT_ALREADY_FAILS"
STABLE_PATCH_FAIL = "STABLE_PATCH_FAIL"
NO_FAILURE = "NO_FAILURE"
ENVIRONMENT_SKEW = "ENVIRONMENT_SKEW"
INCOMPLETE = "INCOMPLETE"
@dataclass(frozen=True)
class Run:
runner_id: str
tree: str
passed: bool
def cell_of(flags: list[bool]) -> Cell:
if not flags:
return Cell.EMPTY
if all(flags):
return Cell.STABLE_PASS
if not any(flags):
return Cell.STABLE_FAIL
return Cell.MIXED
def classify(
runs: list[Run],
k: int,
runners: tuple[str, str],
digests: dict[str, str] | None = None,
) -> Decision:
if k < 2:
return Decision.INCOMPLETE
if digests is not None:
seen = [digests.get(r) for r in runners]
if not all(seen) or len(set(seen)) != 1:
return Decision.ENVIRONMENT_SKEW
allowed = set(runners)
buckets: dict[tuple[str, str], list[bool]] = defaultdict(list)
for run in runs:
if run.runner_id not in allowed or run.tree not in {"patch", "revert"}:
return Decision.INCOMPLETE
buckets[(run.runner_id, run.tree)].append(run.passed)
needed = [(r, t) for r in runners for t in ("patch", "revert")]
if any(len(buckets[key]) != k for key in needed):
return Decision.INCOMPLETE
cells = {key: cell_of(buckets[key]) for key in needed}
patch_cells = {cells[(r, "patch")] for r in runners}
revert_cells = {cells[(r, "revert")] for r in runners}
if len(patch_cells) != 1 or len(revert_cells) != 1:
return Decision.ENVIRONMENT_SKEW
patch_cell = patch_cells.pop()
revert_cell = revert_cells.pop()
if patch_cell is Cell.STABLE_PASS:
return Decision.NO_FAILURE
if patch_cell is Cell.STABLE_FAIL:
return Decision.STABLE_PATCH_FAIL
if revert_cell is Cell.MIXED:
return Decision.HARNESS_NOT_PATCH
if revert_cell is Cell.STABLE_FAIL:
return Decision.PARENT_ALREADY_FAILS
if revert_cell is Cell.STABLE_PASS:
return Decision.PATCH_FREEZE_CANDIDATE
return Decision.INCOMPLETE
def demo() -> None:
runners = ("runner-a", "runner-b")
def block(tree: str, pattern: list[bool]) -> list[Run]:
return [Run(runner, tree, flag) for runner in runners for flag in pattern]
cases = {
"patch owns the flake": block("patch", [True, False]) + block("revert", [True, True]),
"harness owns the flake": block("patch", [True, False]) + block("revert", [True, False]),
"parent already fails": block("patch", [True, False]) + block("revert", [False, False]),
"stable patch miss": block("patch", [False, False]) + block("revert", [True, True]),
}
for name, runs in cases.items():
print(name, classify(runs, k=2, runners=runners).value)
if __name__ == "__main__":
demo()
Proposed checks, not a recorded suite run:
def test_digest_mismatch_is_skew():
runs = []
for runner in ("runner-a", "runner-b"):
runs += [Run(runner, "patch", True), Run(runner, "patch", False)]
runs += [Run(runner, "revert", True), Run(runner, "revert", True)]
digests = {"runner-a": "aaa", "runner-b": "bbb"}
assert classify(runs, 2, ("runner-a", "runner-b"), digests) is Decision.ENVIRONMENT_SKEW
Run the demo with python witness_freeze.py. Expected lines are the four decision names in the cases dict, in insertion order.
Numbered workflow
Pick k first and write it into the bundle. Adaptive reruns until green are not a sample. Use k >= 2. Larger k costs runner time and reduces the chance that a rare miss looks stable. This article does not set a time budget.
- Capture seed, raw input, and fixture digest from the first red property check. Do not edit fixtures after that capture. Do not replace
witness_idwith a shortened input discovered later. - If the agent patch is the tip commit, hash
git show --binary --format= HEADforpatch_id, and setrevert_idfromgit rev-parse 'HEAD^^{tree}'. If the patch is still uncommitted, hashgit diff --binary HEADinstead, and usegit rev-parse 'HEAD^{tree}'as the revert tree. - On runner A, execute the property
ktimes on the patch tree andktimes on the revert tree. Store booleans. Keep logs as attachments. Do not feed log text into the classifier. - Repeat step 3 on runner B from a clean checkout. No reused volume and no copied
.pytest_cache. A free server seat is optional here. Isolation is the requirement. - Refuse the bundle when any cell has fewer than
kruns, when fixture digests differ, or when an unknown runner id appears. - Call
classify. Write a patch-scoped freeze only forPATCH_FREEZE_CANDIDATE. Persist every other enum value as a refusal so a later audit can see the denial reason. - On
HARNESS_NOT_PATCH, open the ticket against the fixture or the property harness. Do not hang the waiver on the agent patch id.
Commands that keep the ids tied to bytes:
git show --binary --format= HEAD | sha256sum | awk '{print $1}'
git rev-parse 'HEAD^^{tree}'
find fixtures -type f -print0 | sort -z | xargs -0 sha256sum | sha256sum
The third command hashes a listing of path plus content hash. It misses empty directories and file modes. If mode bits affect the property, add find -printf mode data to the listing before the outer hash. Hash the bytes you replay, not a prose summary of the diff.
Read the matrix as a design constraint
Two trees, three labels each, and a requirement that runners agree: nine pairs. Only one pair is a patch-scoped freeze. If runners may disagree, each tree has nine runner-pair states and the product is 81. Those counts describe the policy. They are not a field rate for agent patches.
| Patch class, runners agree | Revert class, runners agree | Decision |
|---|---|---|
| MIXED | STABLE_PASS | Patch-scoped freeze candidate |
| MIXED | MIXED | Harness, not the patch |
| MIXED | STABLE_FAIL | Parent already fails |
| STABLE_FAIL | any | Stable miss, no freeze |
| STABLE_PASS | any | Nothing to freeze |
Disagreement collapses the row to ENVIRONMENT_SKEW. Averaging runners would hide the isolation failure that runner B was added to catch.
Store refusals in the same shape as freezes
Keep a JSON record next to the patch. A comment in the test file is not a record. Use the same schema for refusals and for the single candidate decision, so deletion of a red log line is visible as a missing file rather than as a quiet pass.
{
"decision": "HARNESS_NOT_PATCH",
"witness_id": "9c1e",
"patch_id": "b7aa",
"revert_id": "0d44",
"k": 2,
"runners": ["runner-a", "runner-b"],
"cells": {
"patch": "MIXED",
"revert": "MIXED"
}
}
The identifiers above are placeholders for the schema. Replace them with the digests from your own commands. Do not copy them into a real freeze file.
Limitations
The revert control needs a diff you can remove. A squashed blob that mixes formatting, lockfile churn, and logic does not have a meaningful revert cell. Split that change, or refuse the freeze.
Shared downstream state fools both runners together. Two clean checkouts that call one staging API can both land on MIXED because the API is mixed. The checker will blame the harness. That may be right, and the harness may sit outside the repository. Confirm the dependency before you file the ticket.
k = 2 can still label a rare miss as stable by chance. Raise k only as far as the runners you actually have. Do not invent extra repeats you will not execute.
A property that holds on both trees can still be a weak property. This gate does not rank assertion strength. Pair it with a separate review of what the predicate excludes if that is the risk you care about.
Fixture listing hashes drift when setup writes timestamps into files under fixtures. Canonicalize or exclude those paths before hashing, or every run becomes environment skew and the matrix never closes.
Who should skip this
Skip the workflow if you cannot build a revert tree, if runner B cannot be cleaned between attempts, or if the debate is about an exploratory session rather than an executable property. Also skip it when a team would accept a model paragraph as a substitute for the boolean matrix. The matrix is the decision. The paragraph is optional drafting.
If a free server seat is already how you obtain a clean second checkout, bind that seat to runner B and leave classify in the repository. The freeze file should cite witness_id and the decision enum, not a vendor name.
Top comments (0)