DEV Community

Finley Zhou
Finley Zhou

Posted on

A Looser Bound Is Not Eligible for a Flake Freeze

A flake freeze covers noise. It does not cover a patch that loosens a numeric bound, drops an assertion id, or shrinks the fixture set. Those edits are deterministic weakenings, and a mixed CI run does not change their class. This workflow classifies the predicate diff first, then spends a small replay budget only if the predicate and the fixture hash are unchanged.

Agent patches make the split easy to miss. One diff can edit production code and the oracle together, the suite goes green or flaps, and a freeze request looks like hygiene. The bound moved, so a freeze written then records a pardon under a flake label.

The classifier below does not merge anything. It answers a narrower question: may this miss even ask for a freeze?

What still counts as a flake

A flake, here, is a narrow event. The property spec is stable under a canonical hash. The fixture-set hash matches the base. A fixed replay budget returns both passes and fails, and the fail count stays inside a stated budget.

A looser bound is the opposite event. The candidate raises max, lowers min, or deletes a required assertion id. That change can explain a green or mixed result with no nondeterminism at all. Fixture shrinkage belongs in the same class. Four cases became two. The missing cases are absent, not flaky.

Tightening is not a free pass in the other direction. A lower max can sit under a real service floor. The script reports bound_tightened, refuses the freeze, and still does not label the patch as a proven regression. That row needs the owner of the metric.

Decision table

Treat the table as policy, not as a summary of a measured corpus. Thresholds are parameters. Replace them with the replay budget your service owner already accepts, and do not tune them from a single red job.

Predicate diff Fixture hash Replay pattern Freeze Next action
Bound loosened, assertion removed, or fixture ids removed any any ineligible Reject the test edit. Do not write a freeze.
Bound tightened, or fixture ids only added any any ineligible Not a flake. Domain owner reviews. Do not auto-reject as a loosening.
Unchanged changed any ineligible Freeze cannot attach. A human reviews the new fixture.
Unchanged unchanged all pass not requested No freeze. Other gates may still run.
Unchanged unchanged all fail ineligible Stable regression. Reject the patch.
Unchanged unchanged mixed, fails at or under budget request only A human may write a freeze keyed to both hashes.
Unchanged unchanged mixed, fails over budget ineligible Harness is too noisy. Fix it before any freeze label.
Unchanged unchanged empty or unknown tokens ineligible Replay file is invalid. Do not guess.

1. Lock a spec the diff can see

Put the property in a small JSON file beside the fixture. The classifier reads that file. It does not parse arbitrary test source. If your checks live only in free-form asserts, this gate will not see them, and that blind spot is a limitation rather than a silent pass.

{
  "property_id": "latency_budget",
  "min": 0,
  "max": 40,
  "assertions": ["status_ok", "no_partial_write"],
  "fixture_ids": ["case_01", "case_02", "case_03", "case_04"]
}
Enter fullscreen mode Exit fullscreen mode

The numbers are illustrations. A max of 40 is not a recommended latency budget. Hash the canonical JSON, with sorted keys and no insignificant whitespace, so a reformat cannot pretend the predicate moved. If a human later writes a freeze, store the full hash and the fixture-set hash on the record. A later patch that changes either hash cannot reuse it.

2. Diff the base spec against the candidate

Run the diff before any replay and before any model call. A bound change is enough to stop. The replay file is optional for weakening checks. Omit it, and a looser bound still fails closed.

python bound_freeze_gate.py \
  --base specs/latency_budget.base.json \
  --candidate specs/latency_budget.head.json \
  --fail-budget 2
Enter fullscreen mode Exit fullscreen mode

Materialize base from the merge-base spec. Materialize candidate from the patch, including a spec drafted by a model. Both are files. Prose about why the bound should move is not an input.

3. Classify the replay only after the diff is empty

Spend a fixed local budget once the predicate is clear. This example uses five runs and a fail budget of two. Those values are knobs. They are not a measured sweet spot.

{
  "fixture_hash": "fixture-set-hash-from-your-store",
  "base_fixture_hash": "fixture-set-hash-from-your-store",
  "results": ["pass", "pass", "fail", "pass", "pass"]
}
Enter fullscreen mode Exit fullscreen mode
python bound_freeze_gate.py \
  --base specs/latency_budget.base.json \
  --candidate specs/latency_budget.head.json \
  --replay replays/latency_budget.json \
  --fail-budget 2
Enter fullscreen mode Exit fullscreen mode

All fail means a stable miss. Mixed with more fails than the budget means the harness is not ready for a freeze label. Mixed inside the budget is the only row that may emit freeze_request. Empty results, or tokens other than pass and fail, are invalid. The script must not treat them as green.

4. Keep draft models off the admit path

Disclosure: This article was prepared as part of MonkeyCode's product outreach. MonkeyCode's free model access fits one step earlier, and only as a draft. Use it to propose a candidate spec from a failing trace, or to draft the production patch the spec will judge. Write that proposal to disk. It then enters the same diff as a human edit.

A fluent argument for raising max from 40 to 80 does not change the row in the table. The free server option fits the replay step, when you want a clean fixture directory away from leftover processes on a laptop. A local container with a fresh worktree is an equivalent choice. If free model access is unavailable, or the completion comes back empty, skip the draft and keep the current spec. If the free server is unavailable, run the same command locally.

Neither outage is a reason to skip the diff. Neither success is a reason to admit the patch. Do not send the freeze decision back to a model for a second opinion. The decision is the table. If you already use that free model access to draft patches, put this command between the draft file and any freeze write.

5. Run the reference classifier

The script is a proposal you can execute on the JSON files above. It is not a hosted service. This article does not claim a detection rate, a corpus score, or a production incident count for it.

#!/usr/bin/env python3
"""Decide whether an agent-patch miss may request a flake freeze.

Proposal only. The exit code is a review signal, not a merge.
"""

import argparse
import hashlib
import json
import sys

def canonical_hash(payload):
    raw = json.dumps(payload, sort_keys=True, separators=(",", ":"))
    return hashlib.sha256(raw.encode("utf-8")).hexdigest()

def predicate_findings(base, candidate):
    findings = []
    if base.get("property_id") != candidate.get("property_id"):
        findings.append("property_id_changed")
    if candidate["max"] > base["max"] or candidate["min"] < base["min"]:
        findings.append("bound_loosened")
    if candidate["max"] < base["max"] or candidate["min"] > base["min"]:
        findings.append("bound_tightened")
    removed = sorted(set(base["assertions"]) - set(candidate["assertions"]))
    if removed:
        findings.append("assertion_removed")
    if set(base["fixture_ids"]) != set(candidate["fixture_ids"]):
        findings.append("fixture_set_changed")
        dropped = sorted(set(base["fixture_ids"]) - set(candidate["fixture_ids"]))
        if dropped:
            findings.append("fixture_shrunk")
    return findings

def replay_counts(replay):
    results = replay.get("results")
    if not isinstance(results, list) or not results:
        return None
    if any(item not in ("pass", "fail") for item in results):
        return None
    fails = sum(item == "fail" for item in results)
    passes = sum(item == "pass" for item in results)
    return fails, passes, len(results)

def decide(findings, replay, fail_budget):
    hard_reject = {
        "bound_loosened",
        "assertion_removed",
        "fixture_shrunk",
        "property_id_changed",
    }
    if hard_reject.intersection(findings):
        return "reject_weakening", "freeze_ineligible"
    if "bound_tightened" in findings or "fixture_set_changed" in findings:
        return "predicate_changed", "freeze_ineligible"
    if replay is None:
        return "predicate_clear", "replay_required"
    if replay.get("fixture_hash") != replay.get("base_fixture_hash"):
        return "fixture_moved", "freeze_ineligible"
    counted = replay_counts(replay)
    if counted is None:
        return "replay_invalid", "freeze_ineligible"
    fails, passes, total = counted
    if fails == 0 and passes == total:
        return "stable_pass", "no_freeze"
    if passes == 0 and fails == total:
        return "stable_fail", "freeze_ineligible"
    if fails <= fail_budget:
        return "mixed_within_budget", "freeze_request"
    return "mixed_over_budget", "freeze_ineligible"

def main():
    parser = argparse.ArgumentParser()
    parser.add_argument("--base", required=True)
    parser.add_argument("--candidate", required=True)
    parser.add_argument("--replay")
    parser.add_argument("--fail-budget", type=int, default=2)
    args = parser.parse_args()
    if args.fail_budget < 0:
        parser.error("--fail-budget must be >= 0")

    with open(args.base, encoding="utf-8") as handle:
        base = json.load(handle)
    with open(args.candidate, encoding="utf-8") as handle:
        candidate = json.load(handle)
    replay = None
    if args.replay:
        with open(args.replay, encoding="utf-8") as handle:
            replay = json.load(handle)

    findings = predicate_findings(base, candidate)
    kind, freeze = decide(findings, replay, args.fail_budget)
    report = {
        "property_id": base.get("property_id"),
        "base_hash": canonical_hash(base),
        "candidate_hash": canonical_hash(candidate),
        "findings": findings,
        "kind": kind,
        "freeze": freeze,
    }
    json.dump(report, sys.stdout, indent=2)
    sys.stdout.write("\n")
    if freeze == "no_freeze":
        return 0
    if freeze == "freeze_request":
        return 2
    return 1

if __name__ == "__main__":
    sys.exit(main())
Enter fullscreen mode Exit fullscreen mode

Exit codes are part of the contract. Zero means the replay was stably green, so no freeze should be written. Two means a human may open a freeze request keyed to both hashes. One means stop: do not freeze, and do not auto-merge.

A wrapper that treats every non-zero exit as "flake, ignore" undoes the gate. Map code 2 to a labeled review job, not to a skip. Read the JSON kind when you need to separate a weakening reject from a tightened bound that only needs an owner. An uncaught crash is also a stop. Do not wrap the script in a shell clause that turns tracebacks into green.

Expected classifications, given the rules above, are mechanical. A candidate whose max is 80, with the same assertion ids, produces bound_loosened, reject_weakening, and freeze_ineligible. The same spec on both sides, with the five-run file shown earlier and a fail budget of two, produces mixed_within_budget and freeze_request. Neither line is a measured flake rate. Both are outputs of this function.

Added assertion ids are ignored on purpose. Presence of a new id is not proof the new check is correct. A deleted id is the event this gate is built to catch.

6. Keep the patch job in one order

The order is the control. Reordering the steps recreates the failure this classifier exists to catch.

  1. Write base from the merge-base spec and candidate from the patch.
  2. Run the classifier with no replay file. If findings include a weakening, stop. Do not spend a model call to justify the bound.
  3. If the predicate is clear, run the fixture replay on a clean machine. Use a fresh worktree, whether that is the free server option or a local container.
  4. Pass the replay JSON through the same script.
  5. On freeze_request, open a review item that stores property_id, base_hash, candidate_hash, fixture_hash, and the replay token list. Leave the reason field as those hashes, not a model transcript.
  6. On stable_fail, file a regression. A freeze record is the wrong artifact.

A freeze_request item is still not a freeze. It becomes one only after a person confirms the miss is environmental. A record missing either hash cannot be matched later, so it must not be stored. Omit the model transcript from that record.

What the script cannot see

The schema misses a weakened check that exists only in Python. An agent can leave the JSON untouched and delete the assert that enforced it. Generate the assert from the spec, or lock the test file with a separate textual diff, so the JSON stays the source of truth. This script does not generate asserts.

A mixed replay can be a dirty environment rather than a product flake. Eligibility is not permission to merge. Teams that auto-merge on freeze_request should not adopt these exit codes as written.

Skip the workflow if you have no fixture store, if properties are comments, or if the same model call both edits the bound and is allowed to approve the edit. Also skip it when you need a general quarantine for legacy UI tests. This gate is for agent patches that can rewrite the oracle in the same diff as the fix.

Store the JSON report beside the patch. The next reviewer should rerun the command and match both hashes. If the hashes diverge, the report is stale. Discard it.

Top comments (0)