DEV Community

Finley Zhou
Finley Zhou

Posted on

Green Property Checks Must Not Refill a Flake Freeze Token

A flake freeze is a single-spend token. A later property pass does not refill its replay budget, and a fixture match does not move its expiry. Those three results answer different questions. Folding them into one green bit is how a temporary waiver becomes a permanent merge rule.

Agent patches make that collapse easy. A diff can preserve every locked output and still break an invariant the fixture never named. A diff can trip one known flake and still be the wrong change to merge. Status codes hide the split. This note specifies a write policy for the freeze ledger: automated checks may spend a token, and they may not mint or refill one.

The program below is an illustrative specification with a fixed clock of 2026-10-09T00:00:00Z. It is not a report of production failure rates. The outcomes listed with it are the ones the example is written to produce.

Keep three answers in three fields

Property checks state whether the diff preserves invariants you can evaluate without the golden file. Fixture locks state whether observed outputs still hash to the locked digest. A freeze token states whether a reviewer already classified this failure signature as non-blocking, and whether that classification still has budget.

Each check writes only its own field. A property pass stores property=pass. It does not increment remaining_replays. A fixture match stores fixture=match. It does not assign a new expires_at. Only a mint command creates a token, and that command requires a reviewer id the patch author cannot supply.

Abstain is not a pass. If the invariant module fails to import, the verdict is abstain and admission stops. If the fixture file is absent, the verdict is missing, not match. Lookup order is not the claim here. The claim is which fields a check may write once a token row is already in hand.

Read the matrix before you automate it

Rows are inputs. The last column is the only allowed action. write is true only when the ledger row itself must change.

Property Fixture Token Action Write row
fail any any reject, do not read the token no
abstain any any quarantine, do not mint no
pass missing any quarantine until a fixture lock exists no
pass mismatch none or expired reject, no inherited budget only to mark expiry if a row exists and is spent out
pass mismatch live, budget > 0 spend one, hold, do not merge yes, budget only
pass match any admit, leave the token untouched no

The last row is the non-refill rule. A match may admit this patch. It must not extend some other token that shares the test id. Two signatures are different rows. Hash the test id, the normalized message, and the fixture id. Do not put the patch body in that key.

A regenerated diff is a new patch_id, but it can still spend a live token if it reproduces the same signature. It still cannot refill that token. Spending and refilling are different writes. Keeping them apart is the whole gate.

Normalize the signature, then mint

Unstable text defeats the budget. Timestamps and random ids inside assertion messages create a new key on every run, so the ledger never spends the row you meant to cap.

  1. Strip timestamps and generated ids from the failure message before hashing.
  2. Build failure_signature from test id, normalized message, and fixture id.
  3. Reject the mint if any of those parts is empty.
  4. Store reviewer, reason_code, remaining_replays, and expires_at in UTC on the ledger clock.
  5. Refuse mint from the process that wrote the patch. A separate command must supply the reviewer id.

A local normalization check, using a fixed sample rather than a live log:

python3 - <<'PY'
import hashlib, re
msg = "race on list at 2026-10-09T11:02:03Z id=91ab"
norm = re.sub(r"\d{4}-\d{2}-\d{2}T[\d:.]+Z", "<ts>", msg)
norm = re.sub(r"id=[0-9a-f]+", "id=<id>", norm)
print(norm)
print(hashlib.sha256(b"t_order|" + norm.encode() + b"|fx_order").hexdigest()[:16])
PY
Enter fullscreen mode Exit fullscreen mode

The specified normalized message is race on list at <ts> id=<id>. If your real logs do not reduce to a stable string under the same substitutions, do not mint. Add a substitution, or drop the freeze. A token on a moving key is an unbounded waiver with a hash in front of it.

Apply the write policy in order

  1. Hash the diff and store patch_id. Do not call a model in this step.
  2. Run property checks in-process. Record pass, fail, or abstain. On fail or abstain, stop. Do not open the token store.
  3. Replay locked fixtures. Record match, mismatch, or missing. On missing, stop and open a lock task.
  4. On mismatch, derive failure_signature and load the token for that signature only.
  5. If the row is missing, or expires_at is not after the ledger clock, reject. Do not auto-mint.
  6. If remaining_replays is zero, mark the row expired and reject. That write records exhaustion. It does not grant a new budget.
  7. If budget remains, decrement it by one with a compare-and-swap, then hold. A hold is not a merge.
  8. On pass plus match, admit and set write=false. Do not persist the token object, even when every field looks unchanged.
  9. Mint only through the separate command above. Automation may spend. It may not create, and it may not refill after a pass.

Step 8 is the one teams skip when a green suite feels like evidence that the flake is gone. Green evidence can admit this patch. It is not evidence that a future miss should receive a larger budget. Persisting the returned object when write is false is a defect, because a later edit inside decide could refill fields and the save would commit them.

Specification you can run

Save the following as freeze_token_gate.py. It uses a fixed clock so the result does not depend on the runner's time. It does not open a socket. Treat it as a specification of the write bit, not as a measured CI report.

import hashlib
import json
from datetime import datetime, timedelta, timezone

FIXED_NOW = datetime(2026, 10, 9, tzinfo=timezone.utc)

def signature(test_id, message, fixture_id):
    raw = f"{test_id}|{message.strip()}|{fixture_id}".encode()
    return hashlib.sha256(raw).hexdigest()[:16]

def mint(test_id, message, fixture_id, reviewer, budget, hours, now=FIXED_NOW):
    if not reviewer or budget < 1 or hours < 1:
        raise ValueError("mint requires reviewer, budget >= 1, hours >= 1")
    return {
        "signature": signature(test_id, message, fixture_id),
        "reviewer": reviewer,
        "reason_code": "known-order-flake",
        "remaining_replays": budget,
        "expires_at": now + timedelta(hours=hours),
    }

def decide(prop, fixture, token, now=FIXED_NOW):
    if prop == "fail":
        return "reject", token, False
    if prop == "abstain" or fixture == "missing":
        return "quarantine", token, False
    if prop == "pass" and fixture == "match":
        return "admit", token, False
    if fixture != "mismatch" or token is None:
        return "reject", token, False
    if token["expires_at"] <= now or token["remaining_replays"] <= 0:
        expired = dict(token)
        expired["remaining_replays"] = 0
        return "reject", expired, True
    spent = dict(token)
    spent["remaining_replays"] = token["remaining_replays"] - 1
    return "hold", spent, True

def demo():
    token = mint("t_order", "race on list", "fx_order", "rev-14", 2, 24)
    cases = [
        ("fail", "mismatch", token),
        ("pass", "match", token),
        ("pass", "mismatch", token),
        ("abstain", "match", token),
    ]
    for prop, fixture, tok in cases:
        action, after, write = decide(prop, fixture, tok)
        print(json.dumps({
            "property": prop,
            "fixture": fixture,
            "action": action,
            "write": write,
            "budget_after": None if after is None else after["remaining_replays"],
            "expiry_unchanged": after is not None and after["expires_at"] == token["expires_at"],
        }))

if __name__ == "__main__":
    demo()
Enter fullscreen mode Exit fullscreen mode

Specified results if you run the file as written:

  • fail returns reject, write=false, budget still 2.
  • pass and match return admit, write=false, budget still 2, same expiry.
  • pass and mismatch return hold, write=true, budget 1, same expiry.
  • abstain returns quarantine, write=false, budget still 2.

Pass the spent object back in for a second hold and the next call must reject with budget 0. That sequence is the replay budget. It is not a retry loop, and it is not a reason to call a model. The expiry timestamp stays on the original mint.

Lock the non-refill rule in tests/test_freeze_token_gate.py:

from freeze_token_gate import FIXED_NOW, decide, mint

def test_admit_does_not_refill_or_extend():
    token = mint("t_order", "race on list", "fx_order", "rev-14", 1, 4)
    action, after, write = decide("pass", "match", token, FIXED_NOW)
    assert action == "admit"
    assert write is False
    assert after["remaining_replays"] == 1
    assert after["expires_at"] == token["expires_at"]

def test_spend_does_not_move_expiry_and_second_spend_rejects():
    token = mint("t_order", "race on list", "fx_order", "rev-14", 1, 4)
    action, spent, write = decide("pass", "mismatch", token, FIXED_NOW)
    assert action == "hold"
    assert write is True
    assert spent["remaining_replays"] == 0
    assert spent["expires_at"] == token["expires_at"]
    action2, expired, write2 = decide("pass", "mismatch", spent, FIXED_NOW)
    assert action2 == "reject"
    assert write2 is True
    assert expired["remaining_replays"] == 0
    assert expired["expires_at"] == token["expires_at"]
Enter fullscreen mode Exit fullscreen mode
python3 freeze_token_gate.py
python3 -m pytest -q tests/test_freeze_token_gate.py
Enter fullscreen mode Exit fullscreen mode

If test_admit_does_not_refill_or_extend fails, the admit path is writing the freeze row. Stop and remove that write before you add any network client to the job. A green suite that also refreshes expires_at is not a stronger suite. It is a waiver that grows on success.

Where a free draft is allowed to sit

Steps 1 through 8 are local. A model call is eligible only after the ledger returns reject for a reason code you have already marked regenerable, and only to produce a new diff with a new patch_id. A hold is not eligible. The miss is already classified, and another draft must not spend or refill that row.

Disclosure: This article was prepared as part of MonkeyCode's product outreach.

MonkeyCode's free model access fits as a draft source for that new diff. Its free server option fits as an isolated replay host, so the failure signature is taken from a clean run rather than from leftover state on a shared runner. Neither service is the freeze authority. The mint command still requires a reviewer id that the draft process cannot supply. Those two availability claims are operator-stated options. They are not a quota, a hardware shape, a latency bound, or a promise that the free path remains. If that path is down, run the same ledger on another replay host. Do not skip the token rules because one endpoint is unreachable.

Eligibility for the draft call is a policy list, not a measured win rate:

  • Regenerable reason codes: format, import-order, unused-symbol.
  • Not regenerable: invariant failure, missing fixture, abstain, expired token, exhausted budget.
  • Never on a hold, and never as a substitute for mint.

Keep the property module and the fixture files out of the draft's editable paths. A model that "fixes" a flake by editing the invariant, or by rewriting the golden file, is manufacturing a pass. Treat empty draft output as abstain, not as a signal to retry the same call. Then leave the token row alone.

If your review log can already store a reviewer id apart from the diff author, run the non-refill test on one flaky signature before you attach any draft client.

Limitations

Signature collision is the main hole. A real regression can emit the same normalized message as a known flake. The ledger will then spend the flake token and hold a patch that should have been rejected. Cap the budget low, and require the reason code to name the flake mechanism, not just the assertion text.

Clock skew is the second hole. Compare expires_at to the ledger clock in UTC. A runner clock can make a dead token look live. That is why the example pins FIXED_NOW, and why production code should not trust log timestamps as the authority.

The in-memory store is not safe under two runners. Both can read budget 1 and both can hold unless the decrement is atomic. That bug looks like a flake and is not one. Use a compare-and-swap on (signature, remaining_replays) before you trust a hold.

This gate does not prove the patch is correct. It proves that three answers were not allowed to overwrite each other. Semantic review stays outside the table. A passing property check is only as strong as the invariants you wrote down. If those invariants are restatements of the fixture file, you have one check with two names.

Free-tier drafts add a further limit. The draft can be empty, off-topic, or a test edit that forces a fixture match. None of those outcomes may set write=true on a freeze row. Availability can also change. Build the replay so a missing draft host degrades to "no regeneration," not to "skip the ledger."

Who should skip it

Skip the ledger if you cannot produce a stable signature. A freeze on a key that changes every run is an unbounded waiver with extra columns.

Skip it if the property command and the fixture command are the same process. You would be storing one answer in two fields and calling the split a control. Split the commands first, or do not pretend the table is doing work.

Skip it on an incident hotfix that must ship before anyone can mint. Use the incident path. Do not mint after the fact to launder that hotfix into the freeze table. A backdated token hides the bypass from the same audit this design is meant to support.

Skip it if you want a model to decide that a failure is flaky. That decision is the mint. Putting it on the process that wrote the patch removes the separation the write bit exists to keep.

What to retain for audit

Keep five fields on every hold or reject: patch_id, property verdict, fixture verdict, token signature or none, and the write bit. Those fields are enough to query whether an admit ever persisted expires_at or remaining_replays. If a later change refills budget on the green path, the query fails without a narrative review.

-- illustrative audit, not a vendor-specific dialect
select patch_id, property, fixture, token_signature, write_bit
from patch_gate_log
where action = 'admit' and write_bit = true;
Enter fullscreen mode Exit fullscreen mode

An empty result is the expected steady state for the non-refill rule. Any row that query returns is a gate defect, even if the suite was green that day.

The operational conclusion is narrow. Spend a freeze token once. Expire it on budget or on the ledger clock. Refuse every automated path, including a free draft, that tries to refill it. Property checks and fixture locks keep their own columns, and a green result is not a mint.

Top comments (0)