DEV Community

Finley Zhou
Finley Zhou

Posted on

Freeze a Flake Only After a Clean Runner Reproduces the Same Class

A flake freeze is a recorded exception, not a pass. An agent patch may receive that exception only after a second runner, started from a clean workspace, reproduces the same failure class on the same fixture hash and the same minimized input. One local red bar is not enough. Agreement is the evidence; disagreement closes the freeze lane.

Agent diffs fail in ways a single log line collapses. The property can be false. The fixture can have drifted. The test can be flaky, or the machine that wrote the patch can be dirty. Those cases need different dispositions. A merge gate that treats every red as a freeze candidate will hide deterministic bugs, while a gate that treats every red as a hard blocker will also stall on a known flake the patch never touched.

Decision matrix

The matrix below is a policy, not a measured benchmark. Rows are exhaustive for the fields the packet records. No timing data and no pass-rate claim sits behind it.

Local class Clean-runner class Property cohort Freeze lane
flake-shaped same class, same fixture hash pass eligible
flake-shaped pass pass refuse; environment miss
flake-shaped different class any refuse; unstable signature
deterministic any any refuse; fix or revert
any unavailable any refuse; fail closed
flake-shaped same class fail refuse; property owns the bug

Read the last two rows as hard stops. A missing clean run is not a waiver. A failed property cohort is not a flake, even when both runners print a similar stack.

Why runner identity is a field

The authoring workspace is a biased sample. It holds editor swap files, a warm dependency cache, a partial environment file, and whatever the agent wrote while iterating. A rerun on that machine often repeats the bias. Calling the second local rerun confirmation double-counts one environment.

Runner identity has to be coarse enough to audit and fine enough to detect reuse. A useful id is the image digest plus a workspace nonce generated at job start, not a hostname alone. Hostnames collide in autoscaled pools. Nonces that persist across jobs also miss the point of the field.

If the two ids match, the vote is a refuse. The clean runner must also echo the fixture hash it actually loaded. A log that says replayed without that echo is an incomplete run, and incomplete runs take the unavailable row.

Prerequisites, not the thesis

Earlier gates already constrain eligibility: shrink the failing input first, pin the fixture hash, reject a looser bound, and do not mint a new freeze token merely because a property check went green. This article does not reopen those rules. It adds a constraint they do not express. A locally perfect packet can still be an artifact of the authoring machine.

Treat those earlier checks as inputs to this vote. If shrink did not run, do not set local_class to flake_shaped. Leave the class empty and let the checker refuse.

Numbered workflow

  1. Pin identities before any rerun. Record patch_hash, test_id, fixture_hash, minimized_input_hash, and oracle_version. If any field is empty, stop. A freeze without those keys cannot be audited on the next patch.

  2. Classify the local failure into a closed set: flake_shaped, deterministic, or environment. Map assertion-text edits and bound edits into deterministic unless a separate waiver record already exists. Do not add a fourth class inside the runner script to force a green vote.

  3. Run the property cohort on inputs the suspected flake does not own. The cohort must pass on the candidate patch. A pass here does not authorize a freeze. It only shows the patch did not break the unrelated invariant set you chose to lock.

  4. Start a clean runner that did not author the patch and does not reuse the local workspace, dependency cache, or editor scratch files. Replay only the minimized input against the pinned fixture hash. Store the remote class, the remote fixture hash, the remote input hash, and the remote runner id.

  5. Compare, then persist. Eligibility requires equal failure class, equal fixture hash, equal minimized-input hash, a passing property cohort, a completed clean run, and two distinct runner ids. Write packet.json first. A second command consumes the packet and either prints eligible or a refuse reason. Do not fold that decision into a YAML conditional that nobody reviews.

  6. Expire the packet when oracle_version or patch_hash changes. A new patch is a new vote, even when the test name matches a freeze that was valid yesterday.

Property cohort and clean replay may run in parallel. The vote waits for both. A fast property pass must not short-circuit a missing clean run.

Build the runner id in the job, not by hand:

nonce=$(openssl rand -hex 8)
image=$(docker image inspect "$CI_IMAGE" --format '{{.Id}}' 2>/dev/null || echo "local-uncontainerized")
printf 'img:%s+nonce:%s\n' "$image" "$nonce" | tee runner.id
sha256sum fixtures/locked.json cases/minimized.json | tee packet.hashes
git rev-parse HEAD > packet.patch_hash
Enter fullscreen mode Exit fullscreen mode

The local-uncontainerized fallback is honest, not a pass. Two jobs that both print that fallback still need different nonces. If you cannot name the image, say so in the packet and keep the refuse path easy to hit.

Reference checker

The module below is a proposal. This draft does not report an execution against a production suite. The function encodes the vote. It does not discover flakes, shrink inputs, or talk to a network.

from dataclasses import dataclass

ELIGIBLE = "eligible"
REFUSE = "refuse"

@dataclass(frozen=True)
class FreezeVote:
    test_id: str
    patch_hash: str
    fixture_hash: str
    minimized_input_hash: str
    oracle_version: str
    local_runner_id: str
    clean_runner_id: str
    local_class: str
    clean_class: str
    clean_fixture_hash: str
    clean_input_hash: str
    property_cohort_passed: bool
    clean_run_completed: bool

def decide(vote: FreezeVote) -> str:
    required = (
        vote.test_id,
        vote.patch_hash,
        vote.fixture_hash,
        vote.minimized_input_hash,
        vote.oracle_version,
        vote.local_runner_id,
        vote.clean_runner_id,
    )
    if any(not field for field in required):
        return REFUSE
    if not vote.clean_run_completed:
        return REFUSE
    if vote.local_runner_id == vote.clean_runner_id:
        return REFUSE
    if vote.local_class != "flake_shaped":
        return REFUSE
    if vote.clean_class != vote.local_class:
        return REFUSE
    if vote.clean_fixture_hash != vote.fixture_hash:
        return REFUSE
    if vote.clean_input_hash != vote.minimized_input_hash:
        return REFUSE
    if not vote.property_cohort_passed:
        return REFUSE
    return ELIGIBLE
Enter fullscreen mode Exit fullscreen mode

Expected results, if that module is imported as written, are assertions rather than a claimed CI history:

def test_vote_matrix():
    base = dict(
        test_id="prop_balance_holds",
        patch_hash="c0ffee",
        fixture_hash="abc123",
        minimized_input_hash="def456",
        oracle_version="oracle-7",
        local_runner_id="img:sha256:1111+nonce:local",
        clean_runner_id="img:sha256:1111+nonce:clean",
        local_class="flake_shaped",
        clean_class="flake_shaped",
        clean_fixture_hash="abc123",
        clean_input_hash="def456",
        property_cohort_passed=True,
        clean_run_completed=False,
    )
    assert decide(FreezeVote(**base)) == REFUSE
    base["clean_run_completed"] = True
    assert decide(FreezeVote(**base)) == ELIGIBLE
    base["clean_runner_id"] = base["local_runner_id"]
    assert decide(FreezeVote(**base)) == REFUSE
    base["clean_runner_id"] = "img:sha256:1111+nonce:clean"
    base["clean_fixture_hash"] = "drifted"
    assert decide(FreezeVote(**base)) == REFUSE
Enter fullscreen mode Exit fullscreen mode

A synthetic packet shows the on-disk shape. The hashes are placeholders, not an incident report.

{
  "test_id": "prop_balance_holds",
  "patch_hash": "c0ffee",
  "fixture_hash": "abc123",
  "minimized_input_hash": "def456",
  "oracle_version": "oracle-7",
  "local_runner_id": "img:sha256:1111+nonce:local",
  "clean_runner_id": "img:sha256:1111+nonce:clean",
  "local_class": "flake_shaped",
  "clean_class": "flake_shaped",
  "clean_fixture_hash": "abc123",
  "clean_input_hash": "def456",
  "property_cohort_passed": true,
  "clean_run_completed": true
}
Enter fullscreen mode Exit fullscreen mode

Load it only after the hash files exist:

python - <<'PY'
import json, pathlib
from freeze_vote import FreezeVote, decide
raw = json.loads(pathlib.Path("packet.json").read_text())
print(decide(FreezeVote(**raw)))
PY
Enter fullscreen mode Exit fullscreen mode

Suggested process exits, if you wrap decide in a CLI: 0 when the result is eligible, 2 when a required field is empty, 3 when classes or hashes disagree, 4 when the clean run did not complete. Keep the mapping in one module. CI should call the module, not reimplement the table.

Where drafts and a clean workspace fit

Property statements still have to be written, and a model draft is not a verdict. Disclosure: This article was prepared as part of MonkeyCode's product outreach. MonkeyCode's free model access is relevant as a drafting aid: request candidate invariants from the patch diff, then discard any draft that is not executable or that restates the assertion under test. decide never reads a confidence score. It reads hashes, classes, and a boolean the cohort runner computed.

The free server option is relevant as one way to obtain clean_runner_id. Use it as an isolated replay workspace for the minimized case and, if policy allows, the property cohort. Persist the remote fixture hash or refuse the vote. A remote status of success without that hash is still the unavailable row. Nothing here states a quota, a hardware shape, a time limit, or a promise that the option remains free.

Keep packet.json in the repository that owns the merge. A remote log is an input to the vote, not the record of record.

If a second isolated workspace already exists inside the network, point clean_runner_id at that workspace and do not send fixtures anywhere else. The vote depends on distinct runner identity and matching hashes, not on who operates the second machine.

Limitations

Two runners can share a bug. The same base image, the same clock-skew window, or the same external stub will reproduce a false flake twice. Dual agreement raises the bar. It does not prove the failure is nondeterministic in production, and it does not estimate a flake rate.

Non-reproduction is ambiguous. A clean pass means not attested, not that the local failure was imaginary. Retrying until the clean runner fails manufactures eligibility. Cap the clean attempt at one replay of the minimized input. A second try requires a new packet and a written reason field that decide currently ignores on purpose, so a human has to widen the schema knowingly.

Seeded tests only. The policy assumes the minimized input and the fixture hash fully determine the case. UI timing tests, live network calls, and unordered thread dumps do not meet that assumption. Do not coerce them into flake_shaped to make the table close.

The reference checker does not parse pytest, JUnit, or vendor CI logs. Adapter code is separate work. An adapter that drops clean_fixture_hash while still marking clean_run_completed true will hide the gap only if someone later deletes the hash check. Do not delete it to green the pipeline.

Shared image digests are a residual risk, not a reason to skip the nonce. Record both. If you cannot rebuild the clean workspace from the same lockfile as CI, the run is not clean, and the vote should refuse even when ids differ.

Who should not use this

Skip the approach when the case cannot leave the private network and no internal clean runner exists. A shared or public workspace is the wrong place for secrets, customer fixtures, or vulnerability-adjacent inputs. Fail closed instead of redacting logs ad hoc and hoping the class string survived.

Skip it when the suite has no minimized input. Replaying a long full-fixture job on a second machine copies cost and hides the failing bytes. Shrink under the existing rule, or do not freeze.

Skip it when a freeze is expected to mean this red is accepted forever. The packet expires with the patch hash and the oracle version. Teams that need a standing waiver need a different process, with a human owner on the record and a review date. This checker will not invent that owner field for you.

What to merge

Merge when the property cohort passed and the freeze lane was either unused or eligible. Do not merge on a local red the clean runner did not confirm. Do not merge when the clean runner was skipped to save minutes.

Store the packet beside the test for the life of the waiver. The next agent patch should fail this vote until it produces its own attestation, even if the test name matches a previous freeze. That is the point of binding the exception to patch_hash rather than to the test title alone.

Top comments (0)