A flake freeze is a recorded exception, not a pass. An agent patch may receive that exception only after a second runner, started from a clean workspace, reproduces the same failure class on the same fixture hash and the same minimized input. One local red bar is not enough. Agreement is the evidence; disagreement closes the freeze lane.
Agent diffs fail in ways a single log line collapses. The property can be false. The fixture can have drifted. The test can be flaky, or the machine that wrote the patch can be dirty. Those cases need different dispositions. A merge gate that treats every red as a freeze candidate will hide deterministic bugs, while a gate that treats every red as a hard blocker will also stall on a known flake the patch never touched.
Decision matrix
The matrix below is a policy, not a measured benchmark. Rows are exhaustive for the fields the packet records. No timing data and no pass-rate claim sits behind it.
| Local class | Clean-runner class | Property cohort | Freeze lane |
|---|---|---|---|
| flake-shaped | same class, same fixture hash | pass | eligible |
| flake-shaped | pass | pass | refuse; environment miss |
| flake-shaped | different class | any | refuse; unstable signature |
| deterministic | any | any | refuse; fix or revert |
| any | unavailable | any | refuse; fail closed |
| flake-shaped | same class | fail | refuse; property owns the bug |
Read the last two rows as hard stops. A missing clean run is not a waiver. A failed property cohort is not a flake, even when both runners print a similar stack.
Why runner identity is a field
The authoring workspace is a biased sample. It holds editor swap files, a warm dependency cache, a partial environment file, and whatever the agent wrote while iterating. A rerun on that machine often repeats the bias. Calling the second local rerun confirmation double-counts one environment.
Runner identity has to be coarse enough to audit and fine enough to detect reuse. A useful id is the image digest plus a workspace nonce generated at job start, not a hostname alone. Hostnames collide in autoscaled pools. Nonces that persist across jobs also miss the point of the field.
If the two ids match, the vote is a refuse. The clean runner must also echo the fixture hash it actually loaded. A log that says replayed without that echo is an incomplete run, and incomplete runs take the unavailable row.
Prerequisites, not the thesis
Earlier gates already constrain eligibility: shrink the failing input first, pin the fixture hash, reject a looser bound, and do not mint a new freeze token merely because a property check went green. This article does not reopen those rules. It adds a constraint they do not express. A locally perfect packet can still be an artifact of the authoring machine.
Treat those earlier checks as inputs to this vote. If shrink did not run, do not set local_class to flake_shaped. Leave the class empty and let the checker refuse.
Numbered workflow
Pin identities before any rerun. Record
patch_hash,test_id,fixture_hash,minimized_input_hash, andoracle_version. If any field is empty, stop. A freeze without those keys cannot be audited on the next patch.Classify the local failure into a closed set:
flake_shaped,deterministic, orenvironment. Map assertion-text edits and bound edits intodeterministicunless a separate waiver record already exists. Do not add a fourth class inside the runner script to force a green vote.Run the property cohort on inputs the suspected flake does not own. The cohort must pass on the candidate patch. A pass here does not authorize a freeze. It only shows the patch did not break the unrelated invariant set you chose to lock.
Start a clean runner that did not author the patch and does not reuse the local workspace, dependency cache, or editor scratch files. Replay only the minimized input against the pinned fixture hash. Store the remote class, the remote fixture hash, the remote input hash, and the remote runner id.
Compare, then persist. Eligibility requires equal failure class, equal fixture hash, equal minimized-input hash, a passing property cohort, a completed clean run, and two distinct runner ids. Write
packet.jsonfirst. A second command consumes the packet and either printseligibleor a refuse reason. Do not fold that decision into a YAML conditional that nobody reviews.Expire the packet when
oracle_versionorpatch_hashchanges. A new patch is a new vote, even when the test name matches a freeze that was valid yesterday.
Property cohort and clean replay may run in parallel. The vote waits for both. A fast property pass must not short-circuit a missing clean run.
Build the runner id in the job, not by hand:
nonce=$(openssl rand -hex 8)
image=$(docker image inspect "$CI_IMAGE" --format '{{.Id}}' 2>/dev/null || echo "local-uncontainerized")
printf 'img:%s+nonce:%s\n' "$image" "$nonce" | tee runner.id
sha256sum fixtures/locked.json cases/minimized.json | tee packet.hashes
git rev-parse HEAD > packet.patch_hash
The local-uncontainerized fallback is honest, not a pass. Two jobs that both print that fallback still need different nonces. If you cannot name the image, say so in the packet and keep the refuse path easy to hit.
Reference checker
The module below is a proposal. This draft does not report an execution against a production suite. The function encodes the vote. It does not discover flakes, shrink inputs, or talk to a network.
from dataclasses import dataclass
ELIGIBLE = "eligible"
REFUSE = "refuse"
@dataclass(frozen=True)
class FreezeVote:
test_id: str
patch_hash: str
fixture_hash: str
minimized_input_hash: str
oracle_version: str
local_runner_id: str
clean_runner_id: str
local_class: str
clean_class: str
clean_fixture_hash: str
clean_input_hash: str
property_cohort_passed: bool
clean_run_completed: bool
def decide(vote: FreezeVote) -> str:
required = (
vote.test_id,
vote.patch_hash,
vote.fixture_hash,
vote.minimized_input_hash,
vote.oracle_version,
vote.local_runner_id,
vote.clean_runner_id,
)
if any(not field for field in required):
return REFUSE
if not vote.clean_run_completed:
return REFUSE
if vote.local_runner_id == vote.clean_runner_id:
return REFUSE
if vote.local_class != "flake_shaped":
return REFUSE
if vote.clean_class != vote.local_class:
return REFUSE
if vote.clean_fixture_hash != vote.fixture_hash:
return REFUSE
if vote.clean_input_hash != vote.minimized_input_hash:
return REFUSE
if not vote.property_cohort_passed:
return REFUSE
return ELIGIBLE
Expected results, if that module is imported as written, are assertions rather than a claimed CI history:
def test_vote_matrix():
base = dict(
test_id="prop_balance_holds",
patch_hash="c0ffee",
fixture_hash="abc123",
minimized_input_hash="def456",
oracle_version="oracle-7",
local_runner_id="img:sha256:1111+nonce:local",
clean_runner_id="img:sha256:1111+nonce:clean",
local_class="flake_shaped",
clean_class="flake_shaped",
clean_fixture_hash="abc123",
clean_input_hash="def456",
property_cohort_passed=True,
clean_run_completed=False,
)
assert decide(FreezeVote(**base)) == REFUSE
base["clean_run_completed"] = True
assert decide(FreezeVote(**base)) == ELIGIBLE
base["clean_runner_id"] = base["local_runner_id"]
assert decide(FreezeVote(**base)) == REFUSE
base["clean_runner_id"] = "img:sha256:1111+nonce:clean"
base["clean_fixture_hash"] = "drifted"
assert decide(FreezeVote(**base)) == REFUSE
A synthetic packet shows the on-disk shape. The hashes are placeholders, not an incident report.
{
"test_id": "prop_balance_holds",
"patch_hash": "c0ffee",
"fixture_hash": "abc123",
"minimized_input_hash": "def456",
"oracle_version": "oracle-7",
"local_runner_id": "img:sha256:1111+nonce:local",
"clean_runner_id": "img:sha256:1111+nonce:clean",
"local_class": "flake_shaped",
"clean_class": "flake_shaped",
"clean_fixture_hash": "abc123",
"clean_input_hash": "def456",
"property_cohort_passed": true,
"clean_run_completed": true
}
Load it only after the hash files exist:
python - <<'PY'
import json, pathlib
from freeze_vote import FreezeVote, decide
raw = json.loads(pathlib.Path("packet.json").read_text())
print(decide(FreezeVote(**raw)))
PY
Suggested process exits, if you wrap decide in a CLI: 0 when the result is eligible, 2 when a required field is empty, 3 when classes or hashes disagree, 4 when the clean run did not complete. Keep the mapping in one module. CI should call the module, not reimplement the table.
Where drafts and a clean workspace fit
Property statements still have to be written, and a model draft is not a verdict. Disclosure: This article was prepared as part of MonkeyCode's product outreach. MonkeyCode's free model access is relevant as a drafting aid: request candidate invariants from the patch diff, then discard any draft that is not executable or that restates the assertion under test. decide never reads a confidence score. It reads hashes, classes, and a boolean the cohort runner computed.
The free server option is relevant as one way to obtain clean_runner_id. Use it as an isolated replay workspace for the minimized case and, if policy allows, the property cohort. Persist the remote fixture hash or refuse the vote. A remote status of success without that hash is still the unavailable row. Nothing here states a quota, a hardware shape, a time limit, or a promise that the option remains free.
Keep packet.json in the repository that owns the merge. A remote log is an input to the vote, not the record of record.
If a second isolated workspace already exists inside the network, point clean_runner_id at that workspace and do not send fixtures anywhere else. The vote depends on distinct runner identity and matching hashes, not on who operates the second machine.
Limitations
Two runners can share a bug. The same base image, the same clock-skew window, or the same external stub will reproduce a false flake twice. Dual agreement raises the bar. It does not prove the failure is nondeterministic in production, and it does not estimate a flake rate.
Non-reproduction is ambiguous. A clean pass means not attested, not that the local failure was imaginary. Retrying until the clean runner fails manufactures eligibility. Cap the clean attempt at one replay of the minimized input. A second try requires a new packet and a written reason field that decide currently ignores on purpose, so a human has to widen the schema knowingly.
Seeded tests only. The policy assumes the minimized input and the fixture hash fully determine the case. UI timing tests, live network calls, and unordered thread dumps do not meet that assumption. Do not coerce them into flake_shaped to make the table close.
The reference checker does not parse pytest, JUnit, or vendor CI logs. Adapter code is separate work. An adapter that drops clean_fixture_hash while still marking clean_run_completed true will hide the gap only if someone later deletes the hash check. Do not delete it to green the pipeline.
Shared image digests are a residual risk, not a reason to skip the nonce. Record both. If you cannot rebuild the clean workspace from the same lockfile as CI, the run is not clean, and the vote should refuse even when ids differ.
Who should not use this
Skip the approach when the case cannot leave the private network and no internal clean runner exists. A shared or public workspace is the wrong place for secrets, customer fixtures, or vulnerability-adjacent inputs. Fail closed instead of redacting logs ad hoc and hoping the class string survived.
Skip it when the suite has no minimized input. Replaying a long full-fixture job on a second machine copies cost and hides the failing bytes. Shrink under the existing rule, or do not freeze.
Skip it when a freeze is expected to mean this red is accepted forever. The packet expires with the patch hash and the oracle version. Teams that need a standing waiver need a different process, with a human owner on the record and a review date. This checker will not invent that owner field for you.
What to merge
Merge when the property cohort passed and the freeze lane was either unused or eligible. Do not merge on a local red the clean runner did not confirm. Do not merge when the clean runner was skipped to save minutes.
Store the packet beside the test for the life of the waiver. The next agent patch should fail this vote until it produces its own attestation, even if the test name matches a previous freeze. That is the point of binding the exception to patch_hash rather than to the test title alone.
Top comments (0)