An agent patch should not earn a flaky-test freeze unless the miss is unstable. A stable red assertion is a failure. Property violations and fixture drift are failures too, and none of those three classes can be waived.
That split is the whole strategy. Replay count, freeze storage, and a later model note are useful only after the miss has a class. A freeze record that omits the class is not evidence. It is an unlabeled skip.
What the classifier is for
Agent-generated patches fail tests for reasons that look alike in a CI summary. A red row might be a broken invariant, a fixture the patch rewrote, a new test id, or an assertion that flips across replays. Collapsing those rows into "flake" hides the first three. Blocking every flip, with no class, also stalls review on noise.
The workflow below is a proposal for that classification step. The Python file is unexecuted example code, not a benchmark from this account. Replay counts and freeze lifetimes are policy parameters a team writes down. They are not tuned optima, and they are not product limits.
Six classes, one non-negotiable rule
Name the miss before anyone writes a freeze. The rule that does not bend is small. property_violation, fixture_drift, and stable_assertion_fail are closed.
Only unstable_assertion may become a freeze proposal. That proposal does not count as a pass. Mixed outcomes are necessary for the proposal. They are not sufficient if a property is red or the fixture hash moved.
| Class | What you observed | Freeze | Admission effect |
|---|---|---|---|
property_violation |
Any property predicate is false | Never | Reject |
fixture_drift |
Fixture hash differs from the lock | Never | Reject |
stable_assertion_fail |
Properties green, hash unchanged, every replay failed | Never | Reject |
unknown_test |
Test id is new, renamed, or the replay count is short | Never | Human review |
unstable_assertion |
Properties green, hash unchanged, mixed replay outcomes | Proposal only | Not a pass |
stable_pass |
Properties green, hash unchanged, every replay passed | Not applicable | Eligible for diff review |
Read the table by column, not by hope. A stable failure on a locked fixture is the case teams mislabel most often. It looks like a known-bad test. It is a reproducible miss, so the freeze path must return nothing.
Numbered workflow
Work in this order. Later steps must not repair an earlier class.
- Lock fixtures before the patch is applied. Hash file bytes, not filenames alone. Store the digest beside the suite.
- Run property predicates on the patched tree. Stop on the first false predicate. Do not open freeze storage on that path.
- Replay assertions only if every property passed and every fixture hash matches the lock. Use a declared replay count. This example uses 5. That number is a parameter you set in review policy, then keep stable long enough to compare weeks.
- Classify each test id in local code. Do not ask a model to invent the class from a log paste.
- Write a freeze proposal only for
unstable_assertion. Scope it to one test id, bind it to the fixture hash, set an expiry, and setcounts_as_passto false. - Optionally request a diagnosis draft for rows that are already classified. The draft may explain a row. It may not change the class, the expiry, or the exit code.
The property step exits before any freeze file is read. A green property suite is a precondition for interpreting assertion replays. It is not a score you average with them.
Fixture drift is handled the same way. If the patch rewrites a fixture so a weaker assertion can pass, the hash check fails closed. Updating fixture.lock is a separate human edit, with its own review, not a side effect of a green retry.
Commands that match the sketch
The shell below is an unexecuted layout. Replace paths after you review them. The classifier in the next section reads cases.json, not JUnit. Parsing JUnit into that file is an adapter you own. Leaving that adapter out keeps the policy readable.
find fixtures -type f -name '*.json' -print0 | sort -z | xargs -0 sha256sum > fixture.lock
pytest -q tests/properties --tb=line
status=$?
if [ "$status" -ne 0 ]; then
echo "property_violation: freeze storage not opened" >&2
exit "$status"
fi
mkdir -p reports
for i in 1 2 3 4 5; do
pytest -q tests/assertions --tb=no --junitxml="reports/run-${i}.xml"
done
# Adapter not shown: fold reports/*.xml into cases.json.
python classify_waiver.py \
--cases cases.json \
--min-replays 5 \
--freeze-hours 24 \
--out waiver.json
echo "classifier_exit=$?"
Five JUnit files are raw material. They are not the decision. If a run crashes before producing XML, that test id is unknown_test until the replay set is complete. A short replay list must not be upgraded to a freeze.
Classifier you can read in one file
Save this as classify_waiver.py. It is a proposal you can run after review. It is not a result already executed for this article. Python 3.11 is assumed. Annotations are postponed so the file stays readable.
'''Unexecuted proposal: class-first waiver check for agent-patch reports.'''
from __future__ import annotations
import argparse
import hashlib
import json
import sys
from dataclasses import dataclass
from datetime import datetime, timedelta, timezone
from enum import Enum
from pathlib import Path
class Klass(str, Enum):
PROPERTY_VIOLATION = 'property_violation'
FIXTURE_DRIFT = 'fixture_drift'
STABLE_ASSERTION_FAIL = 'stable_assertion_fail'
UNKNOWN_TEST = 'unknown_test'
UNSTABLE_ASSERTION = 'unstable_assertion'
STABLE_PASS = 'stable_pass'
CLOSED = {
Klass.PROPERTY_VIOLATION,
Klass.FIXTURE_DRIFT,
Klass.STABLE_ASSERTION_FAIL,
}
EXIT_BLOCK = CLOSED | {Klass.UNKNOWN_TEST}
@dataclass(frozen=True)
class CaseResult:
test_id: str
property_ok: bool
fixture_hash: str
expected_fixture_hash: str
outcomes: tuple[bool, ...]
known_test: bool
def sha256_file(path: Path) -> str:
'''Hash raw bytes. Must match sha256sum in the lock step.'''
return hashlib.sha256(path.read_bytes()).hexdigest()
def classify(case: CaseResult, min_replays: int = 5) -> Klass:
if min_replays < 1:
raise ValueError('min_replays must be >= 1')
if not case.property_ok:
return Klass.PROPERTY_VIOLATION
if case.fixture_hash != case.expected_fixture_hash:
return Klass.FIXTURE_DRIFT
if not case.known_test or len(case.outcomes) < min_replays:
return Klass.UNKNOWN_TEST
if all(case.outcomes):
return Klass.STABLE_PASS
if any(case.outcomes):
return Klass.UNSTABLE_ASSERTION
return Klass.STABLE_ASSERTION_FAIL
def freeze_proposal(
case: CaseResult,
hours: int = 24,
min_replays: int = 5,
) -> dict | None:
if classify(case, min_replays) is not Klass.UNSTABLE_ASSERTION:
return None
now = datetime.now(timezone.utc)
passed = sum(1 for item in case.outcomes if item)
return {
'test_id': case.test_id,
'class': Klass.UNSTABLE_ASSERTION.value,
'fixture_hash': case.fixture_hash,
'replay_pass': passed,
'replay_total': len(case.outcomes),
'counts_as_pass': False,
'issued_at': now.isoformat(),
'expires_at': (now + timedelta(hours=hours)).isoformat(),
'scope': 'single_test_id',
}
def admission_blocked(cases: list[CaseResult], min_replays: int = 5) -> bool:
return any(classify(case, min_replays) in EXIT_BLOCK for case in cases)
def load_cases(path: Path) -> list[CaseResult]:
raw = json.loads(path.read_text())
cases = []
for row in raw:
row['outcomes'] = tuple(row['outcomes'])
cases.append(CaseResult(**row))
return cases
def main() -> None:
parser = argparse.ArgumentParser(description='Classify agent-patch misses.')
parser.add_argument('--cases', required=True, type=Path)
parser.add_argument('--min-replays', type=int, default=5)
parser.add_argument('--freeze-hours', type=int, default=24)
parser.add_argument('--out', required=True, type=Path)
args = parser.parse_args()
cases = load_cases(args.cases)
rows = []
for case in cases:
klass = classify(case, args.min_replays)
proposal = None
if klass is Klass.UNSTABLE_ASSERTION:
proposal = freeze_proposal(case, args.freeze_hours, args.min_replays)
rows.append({
'test_id': case.test_id,
'class': klass.value,
'freeze': proposal,
})
report = {
'min_replays': args.min_replays,
'blocked': admission_blocked(cases, args.min_replays),
'rows': rows,
}
args.out.write_text(json.dumps(report, indent=2) + '\n')
if report['blocked']:
sys.exit(1)
if __name__ == '__main__':
main()
Two details matter more than the length of the file. freeze_proposal returns None unless the class is exactly unstable_assertion, so a caller cannot attach a freeze to a stable red. admission_blocked ignores proposals on purpose. A queue of exceptions must not flip the gate to green.
The exit set is the three closed classes plus unknown_test. A short replay blocks automation. It still does not earn a freeze. sha256_file is the in-process twin of the lock command. Use one hashing method for the lock and the case file. Mixing sha256sum on bytes with a hash of pretty-printed JSON will look like drift when the fixture did not change.
A specification sketch, not a recorded run
The following checks document the branches. They were not executed for this article. Run them yourself if you adopt the file. By inspection, a mixed row with a green property and a matching hash is the only freeze candidate. The same mix with property_ok false is a property violation. A stable five-failure row is stable_assertion_fail, and it blocks admission.
def _spec_sketch() -> None:
base = dict(
test_id='tests/assertions/test_parse.py::test_round_trip',
property_ok=True,
fixture_hash='abc',
expected_fixture_hash='abc',
known_test=True,
)
stable = CaseResult(**base, outcomes=(False, False, False, False, False))
mixed = CaseResult(**base, outcomes=(False, True, False, True, False))
assert classify(stable) is Klass.STABLE_ASSERTION_FAIL
assert classify(mixed) is Klass.UNSTABLE_ASSERTION
assert freeze_proposal(stable) is None
assert freeze_proposal(mixed)['counts_as_pass'] is False
red_prop = CaseResult(**base, property_ok=False, outcomes=mixed.outcomes)
assert classify(red_prop) is Klass.PROPERTY_VIOLATION
assert admission_blocked([stable, mixed]) is True
That last assert is the point of the matrix. One unstable neighbor does not pardon a stable red. The freeze, if written, sits on the mixed test id alone. The process still exits non-zero because a closed class is present.
A sample cases.json for two rows looks like this. Hashes are placeholders, not digests from a real fixture tree. The second id is a stable fail. The first id is mixed. Together they must block.
[
{
"test_id": "tests/assertions/test_parse.py::test_round_trip",
"property_ok": true,
"fixture_hash": "abc",
"expected_fixture_hash": "abc",
"outcomes": [false, true, false, true, false],
"known_test": true
},
{
"test_id": "tests/assertions/test_parse.py::test_empty",
"property_ok": true,
"fixture_hash": "abc",
"expected_fixture_hash": "abc",
"outcomes": [false, false, false, false, false],
"known_test": true
}
]
After you run the script, waiver.json should show blocked as true. The first row may carry a freeze object. The second row's freeze field should be null. If your output disagrees, treat the script as wrong and stop. Do not hand-edit the class to match a hoped-for merge.
What a legal freeze record contains
Store the proposal next to the report, not inside the test result as a pass. Expiry is required. A record without expires_at is incomplete. A record whose fixture_hash does not match the current lock is void, even if the clock has not run out.
{
"test_id": "tests/assertions/test_parse.py::test_round_trip",
"class": "unstable_assertion",
"fixture_hash": "abc",
"replay_pass": 2,
"replay_total": 5,
"counts_as_pass": false,
"scope": "single_test_id",
"issued_at": "2026-10-09T00:00:00+00:00",
"expires_at": "2026-10-10T00:00:00+00:00"
}
The timestamps above are format examples for a 24-hour policy window, aligned to this draft's review day of 2026-10-09. They are not an observed incident. replay_pass and replay_total keep the instability visible, so a later reader does not reconstruct the mix from memory.
Do not copy one test id's freeze onto a neighbor. Scope stays single_test_id because agent patches often touch several assertions at once. A waiver that covers the file is a blanket skip. When the lock file changes, drop every proposal that still cites the old hash. Do not migrate them forward automatically.
Where a free model and a free server fit
Local classification is the decision. Remote help is optional commentary, and only after the JSON exists.
Disclosure: This article was prepared as part of MonkeyCode's product outreach.
MonkeyCode provides free model access and a free server option. Those two facts are enough to place them in this workflow. They are not enough to promise a model name, a quota, a machine size, a region, or a duration. Re-check availability when you run the job. Do not bake a free tier into a permanent CI assumption.
Use the free server option when the replay loop is the bottleneck: property tests first, then five assertion runs, with JUnit written under reports/. The server does not choose the class. It only executes commands you already reviewed. If that option is busy, withdrawn, or a poor fit for the repo, run the same commands on any other machine. The lock file and the classifier are portable. Admission does not depend on which host produced the XML.
Use free model access only after waiver.json exists. Send one classified row, the fixture hash, and the replay counts. Ask for a short diagnosis of an unstable_assertion: likely sources of instability, and which assertion to read first. Refuse any response that relabels a closed class as a flake. If the call fails, times out, or returns an empty body, keep the local class. The gate does not wait on a narrative, and it does not retry its way into a friendlier label.
A prompt that still respects the matrix is a template, not a transcript from a live call.
You are explaining a classified test row. Do not change its class.
class: unstable_assertion
test_id: tests/assertions/test_parse.py::test_round_trip
fixture_hash: abc
replay_pass: 2
replay_total: 5
counts_as_pass: false
List at most three plausible sources of instability.
Do not recommend a freeze for any other test id.
The prompt is narrow on purpose. A model that marks stable_assertion_fail as flaky has left the policy. Drop that text. The JSON class remains the record a reviewer cites.
Limitations, and who should skip this
This approach assumes properties are real predicates. A predicate that is true for every input will not protect the closed classes. If your property file cannot fail, fix that before you add freeze records. Otherwise the matrix looks strict while the suite cannot see the bug.
The replay parameter is arbitrary until your team documents it. Five runs will miss rare flakes. They will also flag some slow tests as unstable. Raising the count costs time. It does not change the closed set. Do not pick a count because a model suggested it in one session, and do not lower it to clear a queue.
Do not use this gate where any red test is unconditionally forbidden. Safety-critical patches, data-migration scripts, and authentication checks belong in that bucket. A freeze proposal is the wrong tool there, even for a mixed outcome. Policy that says no waivers should delete step 5, not soften step 2.
Do not use the classifier when the same diff edits fixtures and assertions together and expects the lock to follow. That diff is fixture_drift until a reviewer updates fixture.lock on purpose. An agent patch that rewrites the fixture to match a weaker assertion should fail closed. Silent lock updates train the agent to edit the contract instead of the code.
Skip the model step if reviewers start quoting the diagnosis as if it were a test result. The diagnosis is optional prose. waiver.json is the artifact. Also skip the whole freeze path if you cannot store expires_at and counts_as_pass. A boolean mute on a test name will rot, and this matrix will not save it.
What to merge, and what to leave red
Merge review starts only when admission_blocked is false and every remaining non-pass is an unexpired unstable_assertion bound to the current fixture hash. Even then, the freeze is not a pass. The diff still needs a human read. The exception only stops a mixed assertion from being misfiled as either a product defect or a clean green.
Leave stable reds red. Leave property failures red. Leave hash mismatches red. Those three outcomes are the signal that the patch is wrong, or that the fixture contract moved. A freeze store that absorbs them will trend toward silence. Silence is not stability.
If your repository already hashes fixtures, encode these six classes as a pre-merge check and keep counts_as_pass false. Once that file is in review, a free MonkeyCode model can draft a note for a single unstable_assertion row. Keep the note outside the gate.
Top comments (0)