DEV Community

Finley Zhou
Finley Zhou

Posted on

New Tests in an Agent Patch Cannot Inherit a Flake Freeze

A flake freeze applies only to test identities that already existed on the parent commit. If an agent patch adds, renames, or rewrites a test and then points a freeze record at that new identity, the gate is no longer measuring the patch. It is measuring an exemption the patch wrote for itself.

That rule is a set operation, not a judgment call about flake labels. Parent manifest minus the patch's added, renamed, and rewritten test ids is the only eligible set. Anything else stays closed. A green run that depended on a new freeze line does not count.

Eligible rows, and the canary that must stay red

A freeze record is a waiver for one previously observed test. It is not a suite-wide mute. This proposal requires four fields: test_id, parent_sha, assertion_hash, and expires_at. The checker refuses the record when any field is missing, when test_id is absent from the parent manifest, or when expires_at is not in the future relative to the CI clock.

The patch diff is the second input. Added files, rename destinations, and in-place rewrites of test functions are ineligible ids for this change. A freeze that names one of those ids fails the check even if older freeze lines for other ids are valid.

Signal Eligible on this patch Required action
test_id in parent manifest, not touched by the diff, hash matches Yes Keep the existing record
test_id added in this patch No Delete the freeze line and rerun
Rename destination of a parent test No, until a later commit Treat as a new id
Assertion text hash differs from the frozen hash No Do not reuse the old waiver
Canary mutant passes while the ledger is loaded Ledger invalid Block merge

The last row is the proof step. A freeze file can be well-formed and still too wide if the runner applies it to the wrong ids. A checked-in mutant that breaks one property must still fail after the ledger loads.

Illustrative counts, not a measured study: a parent manifest of 12 property tests, a patch that adds 2 and rewrites 1, and a freeze file with 3 records. If one of those records names an added id, check exits 2. The other two records are not consulted as a reason to merge. Partial eligibility is not a pass.

1. Export the parent manifest

Run collection on the parent commit, not on the dirty worktree. The manifest is a list of stable test ids plus a hash of each assertion body. Node ids from a collector are acceptable when they omit timings and local paths.

git checkout --detach "$PARENT_SHA"
python -m pytest tests/properties --collect-only -q > /tmp/parent_collect.txt
python freeze_scope.py manifest --collect /tmp/parent_collect.txt --out parent_manifest.json
git checkout -
Enter fullscreen mode Exit fullscreen mode

If collect-only itself depends on a network call or a model endpoint, stop. A manifest that changes between two collects on the same commit is not a manifest. Fix the collector before anyone discusses flakes.

Store the manifest next to the parent SHA in CI artifacts. Do not regenerate it from the agent branch and call it a baseline. A baseline that already contains the patch cannot detect a self-issued waiver.

2. Classify test ids in the patch

The diff classifier does not decide whether a test is good. It only labels ids as added, renamed_to, rewritten, or untouched.

git diff --name-status "$PARENT_SHA"...HEAD -- tests/properties > /tmp/name_status.txt
git diff -U0 "$PARENT_SHA"...HEAD -- tests/properties > /tmp/test_diff.patch
python freeze_scope.py classify \
  --name-status /tmp/name_status.txt \
  --patch /tmp/test_diff.patch \
  --out patch_tests.json
Enter fullscreen mode Exit fullscreen mode

classify is conservative on purpose. A hunk that changes a function body marks that function rewritten. A rewritten id cannot inherit a freeze in the same patch, even when the file name stayed the same. That blocks a quiet edit that keeps the name and deletes the assertion.

Untouched ids stay eligible only when their assertion hash still matches the parent manifest. A freeze record is not a floating permission attached to a file path. It is attached to one id and one assertion body.

3. Reject self-issued freeze lines

check loads flake_freeze.json and computes the eligible set. Pass an explicit clock so a laptop timezone cannot stretch a record.

python freeze_scope.py check \
  --manifest parent_manifest.json \
  --patch-tests patch_tests.json \
  --freeze flake_freeze.json \
  --now 2026-10-10T00:00:00Z
Enter fullscreen mode Exit fullscreen mode

Exit codes in this proposal:

  • 0 means every freeze line is in the eligible set and unexpired.
  • 2 means a freeze line names an added, renamed, or rewritten id.
  • 3 means an assertion hash mismatch or an unknown parent id.
  • 4 means an expired record was presented as active.

No exit code means the patch is correct. Exit 0 only means the waiver file is not covering tests this change introduced or rewrote. Property failures outside the ledger still fail the build.

4. Apply a canary, then require red

Keep one mutant outside the agent workspace, under mutants/. The mutant should violate a single property the suite claims to enforce. Deleting a lower-bound check in production code is enough. Do not mutate the tests. A mutant that edits tests confuses the scope check with the behavior check.

Apply the mutant in a temporary copy. Then run the suite with the ledger visible to the runner.

python freeze_scope.py apply-mutant \
  --mutant mutants/drop_lower_bound.py \
  --workdir /tmp/canary-tree
FLAKE_FREEZE=/tmp/canary-tree/flake_freeze.json \
  python -m pytest /tmp/canary-tree/tests/properties -q
echo "canary_exit=$?"
Enter fullscreen mode Exit fullscreen mode

Expected result: a non-zero pytest status. If pytest exits 0, the ledger or the runner ignored the property. Merge stops. Do not repair that by adding a freeze line for the canary failure. The canary is not a flake sample.

Repeat the same command in a second clean workspace. Both exit codes must be non-zero. A disagreement is an environment fault. It is not evidence that the property became optional.

5. Reference checker

The module below is a proposal. It has not been executed against a production suite for this note. It does not call a model. Argument parsing for the bash flags above can wrap these functions. The set math is the part that must stay deterministic.

#!/usr/bin/env python3
"""Proposal: scope flake freezes to parent test ids."""

from __future__ import annotations

import hashlib
import json
import sys
from datetime import datetime, timezone
from pathlib import Path

EXIT_OK = 0
EXIT_SELF_FREEZE = 2
EXIT_IDENTITY = 3
EXIT_EXPIRED = 4

def load(path: str) -> dict:
    return json.loads(Path(path).read_text(encoding="utf-8"))

def assertion_hash(text: str) -> str:
    return hashlib.sha256(text.encode("utf-8")).hexdigest()

def eligible_ids(manifest: dict, patch_tests: dict) -> set[str]:
    parent = set(manifest["test_ids"])
    blocked = set(patch_tests.get("added", []))
    blocked.update(patch_tests.get("renamed_to", []))
    blocked.update(patch_tests.get("rewritten", []))
    return parent - blocked

def check(manifest: dict, patch_tests: dict, freeze: dict, now: datetime) -> int:
    allowed = eligible_ids(manifest, patch_tests)
    blocked = set(patch_tests.get("added", []))
    blocked.update(patch_tests.get("renamed_to", []))
    blocked.update(patch_tests.get("rewritten", []))
    for row in freeze.get("records", []):
        test_id = row["test_id"]
        expires = datetime.fromisoformat(row["expires_at"].replace("Z", "+00:00"))
        if expires <= now:
            print(f"expired freeze: {test_id}")
            return EXIT_EXPIRED
        if test_id in blocked or test_id not in manifest["test_ids"]:
            print(f"ineligible freeze: {test_id}")
            return EXIT_SELF_FREEZE
        if test_id not in allowed:
            print(f"unknown freeze: {test_id}")
            return EXIT_IDENTITY
        expected = manifest["assertion_hashes"][test_id]
        if row.get("assertion_hash") != expected:
            print(f"assertion hash mismatch: {test_id}")
            return EXIT_IDENTITY
    return EXIT_OK

def main(argv: list[str]) -> int:
    if len(argv) != 5 or argv[1] != "check":
        print("usage: freeze_scope.py check MANIFEST PATCH FREEZE")
        return EXIT_IDENTITY
    now = datetime.now(timezone.utc)
    code = check(load(argv[2]), load(argv[3]), load(argv[4]), now)
    print(f"exit={code}")
    return code

if __name__ == "__main__":
    sys.exit(main(sys.argv))
Enter fullscreen mode Exit fullscreen mode

assertion_hash is included so a later command can stamp bodies the same way the manifest does. The checker compares stored hashes. It does not decide that a shorter bound is close enough.

A minimal ledger row looks like this:

{
  "records": [
    {
      "test_id": "tests/properties/test_bounds.py::test_lower_bound_holds",
      "parent_sha": "abc123",
      "assertion_hash": "<sha256 of the assertion body on the parent>",
      "expires_at": "2026-10-17T00:00:00Z"
    }
  ]
}
Enter fullscreen mode Exit fullscreen mode

Replace the placeholder hash with the manifest value. A hand-written hash is a mismatch, and mismatch is exit 3. The sample date is a field format, not a recommended waiver length.

Where a draft model and a second workspace fit

Disclosure: This article was prepared as part of MonkeyCode's product outreach.

MonkeyCode's free model access fits one draft task. Use it to propose a candidate mutant, or a candidate property statement, as text. The checker does not read that text as authority. A completion that says a failure is flaky does not create a ledger row. If the endpoint returns an empty body or an unrelated suggestion, discard it and keep the parent manifest. The set difference does not need a model.

The free server option is a second workspace, not a weaker gate. After the local canary fails, run the same mutant and the same ledger there. Both runs must show a non-zero status. A local pass plus a server fail means the laptop environment is dirty. A local fail plus a server pass means the server tree is missing the mutant or the ledger. Neither case authorizes a new freeze line.

Availability is not a quota, a hardware profile, or a permanence claim. Those details are not established in this note. If the free server is down, run the canary twice in two fresh virtualenvs and record both exit codes. Isolation is the requirement. A hosted workspace is only one way to get the second copy.

Limitations

This approach fails open if the parent manifest was generated on a dirty tree. Recompute it from PARENT_SHA only.

Git rename detection misses a test that was deleted and re-added under a new name below the rename threshold. Treat unmatched new ids as added. That will block some legitimate moves. A blocked move is cheaper than a self-issued waiver.

The canary proves one mutant still fails. It does not estimate how many bugs the suite can catch. A broader mutation score needs its own published method. This note does not provide that method, and it does not report a score.

Runtime-generated test ids cannot use this ledger. Stabilize the id first. Skip the approach when the suite has no properties and only snapshot approvals. A snapshot update inside the patch looks like a rewritten test, so the checker will block it. That outcome is appropriate only if snapshots are not your waiver path.

Do not use the checker as a license to keep a failing property green. Expired rows, hash mismatches, and canary passes are hard stops. They are not comments to override in the same commit.

Who should skip it: teams with no parent-commit manifest, suites whose ids embed timestamps, and reviews that need a statistical flake rate rather than a scope gate. This procedure answers a narrower question. Did this patch waive a test it just wrote?

Closing

Compute the eligible set before debating whether a red was flaky. Then load the ledger and require the canary to stay red. If a MonkeyCode free-server workspace is already available, use it as the second copy and store both exit codes in the CI log. The waiver file is valid only when it cannot see the tests this patch just wrote.

Top comments (0)