DEV Community

Finley Zhou
Finley Zhou

Posted on

The Revert Control Decides Whether a Flake Freeze Belongs on the Patch

A patch-scoped flake freeze is valid only when the agent patch owns the intermittency. If the reverted tree is mixed on the same witness, the freeze belongs on the harness, not on the patch. A second green run does not settle that ownership question.

Agent patches fail property checks for reasons that look identical in a CI log. A stable logic miss, a dirty local fixture, and a harness race all print one red line. Waiving every red line hides regressions, and waiving none of them blocks merges on noise. Ownership has to be measured, which is what a revert control is for.

Close the witness bundle first

A witness bundle is a closed record. It is not a pasted traceback. The checker later in this article refuses incomplete records before it applies any freeze policy. Treat the listing as a reference implementation you can run locally. This article does not report a corpus pass rate, a flake frequency, or a production incident count.

Require these fields before classification starts:

  • witness_id: hash of the property name, the seed, and the raw input bytes that failed.
  • fixture_digest: digest of fixture bytes after setup and before the property runs. Include paths in the hashed listing so renamed files do not collide.
  • patch_id: checksum of the agent diff you will replay.
  • revert_id: tree id of the parent, or of that diff with the agent hunks removed.
  • runs: one boolean per planned repeat, tagged with runner_id and tree (patch or revert).
  • k: repeats per cell, chosen before the first rerun. With two runners and two trees, the run list length must be 4 * k.

Runner A is a locked local worktree. Runner B is a clean checkout on another machine or another container. Shared caches, editor swap files, and leftover services disqualify runner A even when the digest string matches.

Use a clean second runner, not a model vote

Disclosure: This article was prepared as part of MonkeyCode's product outreach.

MonkeyCode's free server option is one way to host runner B when the team has no spare clean machine. Free model access can draft a candidate property sentence from the failing input. That sentence is a draft. It counts only after a person turns it into an executable check and commits the code. Model text never occupies a cell in the matrix below. This article states those two availability facts only. It does not name models, quotas, hardware, durations, or permanence.

Drop the product names and the contract is unchanged. Any clean second runner, plus a human-owned executable property, is enough.

Classify cells, then compare trees

Finish every planned run. Do not stop when a cell turns green.

  1. Mark a cell STABLE_PASS when all k runs passed.
  2. Mark it STABLE_FAIL when all k runs failed.
  3. Mark it MIXED when at least one run passed and one failed.
  4. Accept a tree class only when both runners assign that same label.
  5. If the runners disagree, or the fixture digests differ, stop with environment skew. Do not average the booleans into a pass rate.

One pattern allows a patch-scoped freeze. Both runners must mark the patch cell MIXED, and both must mark the revert cell STABLE_PASS. The parent property held. The agent patch is the first tree that became intermittent.

Every other agreement pattern refuses that freeze:

  • Revert is MIXED on both runners. The flake is already in the parent tree, so the target is the harness or the fixture. A waiver attached to patch_id would hide that.
  • Revert is STABLE_FAIL. The parent already misses this witness. The agent change did not create the mode you are about to waive.
  • Patch is STABLE_FAIL. Keep the check closed. This is a miss, not an intermittent label.
  • Patch is STABLE_PASS. There is no failure to freeze.

Reference checker

Save this as witness_freeze.py. It is proposed code for review. Running the demo prints decision names for four constructed matrices. Those prints are not measurements.

from __future__ import annotations

from collections import defaultdict
from dataclasses import dataclass
from enum import Enum

class Cell(str, Enum):
    STABLE_PASS = "STABLE_PASS"
    STABLE_FAIL = "STABLE_FAIL"
    MIXED = "MIXED"
    EMPTY = "EMPTY"

class Decision(str, Enum):
    PATCH_FREEZE_CANDIDATE = "PATCH_FREEZE_CANDIDATE"
    HARNESS_NOT_PATCH = "HARNESS_NOT_PATCH"
    PARENT_ALREADY_FAILS = "PARENT_ALREADY_FAILS"
    STABLE_PATCH_FAIL = "STABLE_PATCH_FAIL"
    NO_FAILURE = "NO_FAILURE"
    ENVIRONMENT_SKEW = "ENVIRONMENT_SKEW"
    INCOMPLETE = "INCOMPLETE"

@dataclass(frozen=True)
class Run:
    runner_id: str
    tree: str
    passed: bool

def cell_of(flags: list[bool]) -> Cell:
    if not flags:
        return Cell.EMPTY
    if all(flags):
        return Cell.STABLE_PASS
    if not any(flags):
        return Cell.STABLE_FAIL
    return Cell.MIXED

def classify(
    runs: list[Run],
    k: int,
    runners: tuple[str, str],
    digests: dict[str, str] | None = None,
) -> Decision:
    if k < 2:
        return Decision.INCOMPLETE
    if digests is not None:
        seen = [digests.get(r) for r in runners]
        if not all(seen) or len(set(seen)) != 1:
            return Decision.ENVIRONMENT_SKEW
    allowed = set(runners)
    buckets: dict[tuple[str, str], list[bool]] = defaultdict(list)
    for run in runs:
        if run.runner_id not in allowed or run.tree not in {"patch", "revert"}:
            return Decision.INCOMPLETE
        buckets[(run.runner_id, run.tree)].append(run.passed)
    needed = [(r, t) for r in runners for t in ("patch", "revert")]
    if any(len(buckets[key]) != k for key in needed):
        return Decision.INCOMPLETE
    cells = {key: cell_of(buckets[key]) for key in needed}
    patch_cells = {cells[(r, "patch")] for r in runners}
    revert_cells = {cells[(r, "revert")] for r in runners}
    if len(patch_cells) != 1 or len(revert_cells) != 1:
        return Decision.ENVIRONMENT_SKEW
    patch_cell = patch_cells.pop()
    revert_cell = revert_cells.pop()
    if patch_cell is Cell.STABLE_PASS:
        return Decision.NO_FAILURE
    if patch_cell is Cell.STABLE_FAIL:
        return Decision.STABLE_PATCH_FAIL
    if revert_cell is Cell.MIXED:
        return Decision.HARNESS_NOT_PATCH
    if revert_cell is Cell.STABLE_FAIL:
        return Decision.PARENT_ALREADY_FAILS
    if revert_cell is Cell.STABLE_PASS:
        return Decision.PATCH_FREEZE_CANDIDATE
    return Decision.INCOMPLETE

def demo() -> None:
    runners = ("runner-a", "runner-b")
    def block(tree: str, pattern: list[bool]) -> list[Run]:
        return [Run(runner, tree, flag) for runner in runners for flag in pattern]
    cases = {
        "patch owns the flake": block("patch", [True, False]) + block("revert", [True, True]),
        "harness owns the flake": block("patch", [True, False]) + block("revert", [True, False]),
        "parent already fails": block("patch", [True, False]) + block("revert", [False, False]),
        "stable patch miss": block("patch", [False, False]) + block("revert", [True, True]),
    }
    for name, runs in cases.items():
        print(name, classify(runs, k=2, runners=runners).value)

if __name__ == "__main__":
    demo()
Enter fullscreen mode Exit fullscreen mode

Proposed checks, not a recorded suite run:

def test_digest_mismatch_is_skew():
    runs = []
    for runner in ("runner-a", "runner-b"):
        runs += [Run(runner, "patch", True), Run(runner, "patch", False)]
        runs += [Run(runner, "revert", True), Run(runner, "revert", True)]
    digests = {"runner-a": "aaa", "runner-b": "bbb"}
    assert classify(runs, 2, ("runner-a", "runner-b"), digests) is Decision.ENVIRONMENT_SKEW
Enter fullscreen mode Exit fullscreen mode

Run the demo with python witness_freeze.py. Expected lines are the four decision names in the cases dict, in insertion order.

Numbered workflow

Pick k first and write it into the bundle. Adaptive reruns until green are not a sample. Use k >= 2. Larger k costs runner time and reduces the chance that a rare miss looks stable. This article does not set a time budget.

  1. Capture seed, raw input, and fixture digest from the first red property check. Do not edit fixtures after that capture. Do not replace witness_id with a shortened input discovered later.
  2. If the agent patch is the tip commit, hash git show --binary --format= HEAD for patch_id, and set revert_id from git rev-parse 'HEAD^^{tree}'. If the patch is still uncommitted, hash git diff --binary HEAD instead, and use git rev-parse 'HEAD^{tree}' as the revert tree.
  3. On runner A, execute the property k times on the patch tree and k times on the revert tree. Store booleans. Keep logs as attachments. Do not feed log text into the classifier.
  4. Repeat step 3 on runner B from a clean checkout. No reused volume and no copied .pytest_cache. A free server seat is optional here. Isolation is the requirement.
  5. Refuse the bundle when any cell has fewer than k runs, when fixture digests differ, or when an unknown runner id appears.
  6. Call classify. Write a patch-scoped freeze only for PATCH_FREEZE_CANDIDATE. Persist every other enum value as a refusal so a later audit can see the denial reason.
  7. On HARNESS_NOT_PATCH, open the ticket against the fixture or the property harness. Do not hang the waiver on the agent patch id.

Commands that keep the ids tied to bytes:

git show --binary --format= HEAD | sha256sum | awk '{print $1}'
git rev-parse 'HEAD^^{tree}'
find fixtures -type f -print0 | sort -z | xargs -0 sha256sum | sha256sum
Enter fullscreen mode Exit fullscreen mode

The third command hashes a listing of path plus content hash. It misses empty directories and file modes. If mode bits affect the property, add find -printf mode data to the listing before the outer hash. Hash the bytes you replay, not a prose summary of the diff.

Read the matrix as a design constraint

Two trees, three labels each, and a requirement that runners agree: nine pairs. Only one pair is a patch-scoped freeze. If runners may disagree, each tree has nine runner-pair states and the product is 81. Those counts describe the policy. They are not a field rate for agent patches.

Patch class, runners agree Revert class, runners agree Decision
MIXED STABLE_PASS Patch-scoped freeze candidate
MIXED MIXED Harness, not the patch
MIXED STABLE_FAIL Parent already fails
STABLE_FAIL any Stable miss, no freeze
STABLE_PASS any Nothing to freeze

Disagreement collapses the row to ENVIRONMENT_SKEW. Averaging runners would hide the isolation failure that runner B was added to catch.

Store refusals in the same shape as freezes

Keep a JSON record next to the patch. A comment in the test file is not a record. Use the same schema for refusals and for the single candidate decision, so deletion of a red log line is visible as a missing file rather than as a quiet pass.

{
  "decision": "HARNESS_NOT_PATCH",
  "witness_id": "9c1e",
  "patch_id": "b7aa",
  "revert_id": "0d44",
  "k": 2,
  "runners": ["runner-a", "runner-b"],
  "cells": {
    "patch": "MIXED",
    "revert": "MIXED"
  }
}
Enter fullscreen mode Exit fullscreen mode

The identifiers above are placeholders for the schema. Replace them with the digests from your own commands. Do not copy them into a real freeze file.

Limitations

The revert control needs a diff you can remove. A squashed blob that mixes formatting, lockfile churn, and logic does not have a meaningful revert cell. Split that change, or refuse the freeze.

Shared downstream state fools both runners together. Two clean checkouts that call one staging API can both land on MIXED because the API is mixed. The checker will blame the harness. That may be right, and the harness may sit outside the repository. Confirm the dependency before you file the ticket.

k = 2 can still label a rare miss as stable by chance. Raise k only as far as the runners you actually have. Do not invent extra repeats you will not execute.

A property that holds on both trees can still be a weak property. This gate does not rank assertion strength. Pair it with a separate review of what the predicate excludes if that is the risk you care about.

Fixture listing hashes drift when setup writes timestamps into files under fixtures. Canonicalize or exclude those paths before hashing, or every run becomes environment skew and the matrix never closes.

Who should skip this

Skip the workflow if you cannot build a revert tree, if runner B cannot be cleaned between attempts, or if the debate is about an exploratory session rather than an executable property. Also skip it when a team would accept a model paragraph as a substitute for the boolean matrix. The matrix is the decision. The paragraph is optional drafting.

If a free server seat is already how you obtain a clean second checkout, bind that seat to runner B and leave classify in the repository. The freeze file should cite witness_id and the decision enum, not a vendor name.

Top comments (0)