DEV Community

Finley Zhou
Finley Zhou

Posted on

An Open Seed Ledger Cannot Authorize a Flake Freeze

An open seed ledger cannot authorize a flake freeze. A property failure on an agent patch is eligible only after the planned seeds were all attempted, the failing seed is named, shrink stayed stable, and the fixture hash still matches the run. Stop early, and the receipt stays open. An open receipt is an incomplete sample, not evidence that the check is flaky.

Agent patches make that distinction easy to miss. The diff is small, the suite is large, and a property runner often exits on the first counterexample. That exit saves minutes. It also erases the seeds that never started, which is the evidence a freeze decision needs.

Suppose the plan is seeds 1101 through 1104, and seed 1102 fails. A runner that exits there has attempted two of four seeds, so the miss rate inside the plan is unknown. The receipt must show an attempted length of 2 and stopped_early set to true. Only a later run that attempts all four can move that same failing seed toward eligibility, and only if shrink is stable.

What a closed ledger proves

A seed ledger is a receipt written before execution, not a tail of whatever the log happened to print. Planned seeds are fixed first. Attempted seeds are appended only when a trial actually starts. If those two lists differ, the run did not finish the declared budget.

Four fields carry the decision: planned seeds, attempted seeds, the failing seed, and a shrink-stable flag. The fixture hash is the fifth field, and it binds the receipt to the inputs the property saw. Change the hash, and you are looking at a different experiment, even when the seed numbers match.

This check does not replace shrinking or fixture identity. It answers a prior question those later reviews assume: did the declared budget finish? Without that answer, silence in the log is being treated as stability.

Receipt shape

Store one JSON receipt next to the patch. The document below is a proposed schema for a local gate. It is not a dump from a production system, and the hash is illustrative rather than computed from a real fixture.

{
  "patch_id": "agent-1842",
  "property_id": "prop_balance_non_negative",
  "planned_seeds": [1101, 1102, 1103, 1104],
  "attempted_seeds": [1101, 1102, 1103, 1104],
  "failing_seed": 1103,
  "shrink_stable": true,
  "fixture_hash": "sha256:9c1e",
  "stopped_early": false,
  "oracle_source": "human_reviewed"
}
Enter fullscreen mode Exit fullscreen mode

oracle_source is required when a model drafts the predicate. Draft text is not a run. It can participate in a close decision only after a person sets human_reviewed and the runner fills seed fields from execution.

Close procedure

  1. Bind the patch id, property id, and fixture hash before any example runs. If the hash is unknown, stop, and do not copy a hash from an older green build.
  2. Write planned_seeds as a fixed list before the first trial. The runner may not append seeds after a failure, and it may not drop seeds to make the receipt look closed.
  3. Execute every planned seed, including seeds after the first failure, until the list is exhausted or the process crashes. A crash sets stopped_early to true and leaves the ledger open.
  4. Record attempted_seeds in execution order, and require the failing seed to be a member of that list. A failure with no seed id is a harness defect, not a flake.
  5. Shrink only the input tied to the failing seed. Set shrink_stable only when two consecutive shrink passes return the same minimal input.
  6. If a model drafted the property text, require oracle_source to equal human_reviewed before close. Otherwise return oracle_unreviewed, even when every planned seed ran.
  7. Replay the closed receipt on a second runner that loads the same fixture hash. A matching failure supports discussing a freeze; a pass on replay means the first failure is still unclassified, so do not delete the failing seed to force agreement.
  8. Store the receipt beside the patch, and retire it when the diff or the fixture hash changes. A same-day calendar stamp is not a reason to keep the freeze.

Decision table

Read the table as a closed set of local status rules. They are conventions for this checker, not a claim about any external test runner's current API.

Planned vs attempted stopped_early shrink_stable oracle_source Result
equal, failing seed present false true human_reviewed freeze eligible
attempted shorter true any any incomplete, no freeze
failing seed absent from attempts false any any harness gap, no freeze
equal false false human_reviewed shrink open, no freeze
equal false true model_draft oracle unreviewed, no freeze
equal, no failing seed false n/a human_reviewed budget pass, no freeze

A budget pass is not a freeze. Freezes exist for named failures that remain after a closed run. Creating one from a green budget inverts the gate: the suite did not fail, so there is nothing to freeze.

Proposed checker

The Python below is an unexecuted proposal. It fails closed when required fields are missing or the plan is empty. It does not shrink inputs, and it does not call a network.

REQUIRED = (
    "planned_seeds",
    "attempted_seeds",
    "stopped_early",
    "shrink_stable",
    "oracle_source",
    "fixture_hash",
)

def close_ledger(receipt: dict) -> str:
    if any(k not in receipt or receipt[k] in (None, "") for k in REQUIRED):
        return "incomplete_fields"
    planned = list(receipt["planned_seeds"])
    attempted = list(receipt["attempted_seeds"])
    if not planned:
        return "empty_plan"
    if receipt["stopped_early"]:
        return "early_stop_incomplete"
    if attempted != planned:
        return "seed_gap"
    failing = receipt.get("failing_seed")
    if failing is None:
        return "budget_pass_no_freeze"
    if failing not in attempted:
        return "seed_mismatch"
    if receipt["oracle_source"] != "human_reviewed":
        return "oracle_unreviewed"
    if not receipt["shrink_stable"]:
        return "shrink_unstable"
    if not str(receipt["fixture_hash"]).startswith("sha256:"):
        return "fixture_unbound"
    return "freeze_eligible"
Enter fullscreen mode Exit fullscreen mode

A negative test should lock the early-stop path. This example is also unexecuted proposal code, meant to sit beside the checker rather than to report a measured CI result.

def test_partial_plan_cannot_freeze():
    status = close_ledger({
        "planned_seeds": [1, 2, 3, 4],
        "attempted_seeds": [1, 2],
        "failing_seed": 2,
        "stopped_early": True,
        "shrink_stable": False,
        "oracle_source": "human_reviewed",
        "fixture_hash": "sha256:abc",
    })
    assert status == "early_stop_incomplete"
Enter fullscreen mode Exit fullscreen mode

Run the checker against a local receipt and the negative test. Any status other than freeze_eligible keeps the freeze path disabled.

python3 close_seed_ledger.py receipts/agent-1842.json
python3 -m pytest tests/test_close_ledger.py -q
Enter fullscreen mode Exit fullscreen mode

Where a free model path and a free server fit

Disclosure: This article was prepared as part of MonkeyCode's product outreach.

Free model access can draft a candidate predicate and a candidate seed list for the receipt. That draft is starting text, not an executed ledger, and this article assigns it no model name, quota, or quality score. The free server option can host the replay in step 7 when the fixture hash is already known. Both points are availability claims only. They are not a hardware specification, a duration promise, or a benchmark result.

The local checker still decides. If the replay process returns a pass while failing_seed is set, keep the receipt open and inspect the fixture load. Do not let a second status overwrite attempted_seeds.

If fixture hashes are already part of your agent-merge gate, add this close step beside that hash and review model-drafted predicates before a receipt is allowed to close.

Limits, and who should skip it

A closed ledger proves the declared seeds ran. It does not prove the oracle is meaningful. A predicate that is true for every input can close a full plan and still test nothing. Review generated assertions separately before you trust a green budget.

Small plans hide rare failures. Closing four seeds does not estimate the failure rate of a much larger input space. Increase planned_seeds explicitly if you need a wider sample. Do not infer a rate from one receipt.

Skip this gate if fixtures change during the run, or if you cannot produce a fixture hash. The hash field would be fiction, and the replay in step 7 would compare different experiments. Also skip it for searches that cannot be given a finite seed list before the first example starts. A ledger cannot close a budget that was never declared.

freeze_eligible is not a merge approval. It only says the failure sample is complete enough to discuss a freeze. The patch can still be wrong, and a deterministic unit test on the same function may still be the better check.

Counts worth keeping

Track three integers per patch: open receipts, closed failures, and budget passes. If open receipts dominate, the runner is stopping early and the freeze path should stay off. If budget passes dominate while the patched function still faults outside CI, the plan is too small or the oracle is weak.

Those integers come from stored receipts. They are not supplied by the draft that proposed the predicate. Keep the three counts next to the patch id, and disable freeze creation whenever the open-receipt count is non-zero for that diff.

Top comments (1)

Collapse
 
suppdevbot profile image
DEV SUPPORTS •

You need to verify your account.

Enter fullscreen mode Exit fullscreen mode

tr.ee/dev-to