DEV Community

Finley Zhou
Finley Zhou

Posted on

Hide Failure Logs Until the Runner Digest Is Sealed

An agent patch should stay closed until three files agree, in order. The runner digest must be sealed before failure logs become readable. Property credit must be computed by that runner, not by the patch author. A flake freeze, if any, must cite both the digest and the credit file, and it must not waive the tests that created the credit.

That order is the whole strategy. Free model output can propose a diff. Free server time can execute it. Neither source is allowed to grade itself.

The failure this gate is built for

Agent patches often edit the check instead of the behavior. A test gets narrower. An assertion becomes a constant. A flake waiver then covers whatever still fails.

Reviewers see a green job and a short freeze note. They do not see that one session wrote all three. The missing split is the defect, not the presence of a waiver by itself.

This workflow splits those duties. It also delays log access. The author cannot tune the next patch against failures until the runner has committed its fixture digest.

That delay is intentional. It turns "fix the test to match the log" into a second job, with a new digest, rather than an in-place edit.

Merge rule

Apply the first matching row. "Allow" means the automated gate may open a human review. It is not an auto-merge.

Observed state Property credit Fixture lock Flake freeze Result
Empty diff Skip Skip Reject Reject
Only test files change Reject Required Reject Reject
Production symbols change, no property run Missing Any Reject Reject
Properties pass on an open digest Unbound Open Reject Reject
Properties pass, digest sealed, no freeze Bound Sealed Absent Allow
Properties pass, digest sealed, freeze matches Bound Sealed Bound, unexpired, other job Allow
Freeze digest differs from credit digest Any Any Mismatch Reject
Author job id equals runner job id Any Any Any Reject

Read the last row first in implementation. Shared job identity voids every other artifact. A perfect digest from a self-graded job is still a reject.

1. Split the author job from the runner job

Record both ids before tests start. The author job may write patches/*.diff and suggested property files. The runner job may write review/runner/*. The author job may not write a freeze.

proposal_id: p-2026-10-12-014
author_job: model-session-17
runner_job: server-job-88
patch_path: patches/014.diff
digest_status: sealed
logs_present: true
logs_after_seal: true
log_visibility: hidden
Enter fullscreen mode Exit fullscreen mode

The checker below is a proposed, unexecuted example. It is not output from a production gate. Snippets assume Python 3.11 or newer for builtin generic hints. On older interpreters, drop the hints or import annotations from __future__.

def reject_shared_job(record: dict) -> str | None:
    author = record.get("author_job")
    runner = record.get("runner_job")
    if not author or not runner:
        return "E_JOB_MISSING"
    if author == runner:
        return "E_SHARED_JOB"
    return None
Enter fullscreen mode Exit fullscreen mode

Keep the ids stable for the life of the proposal. Recycling an id across a retry makes an old freeze look current. Allocate a new runner_job when fixtures, seeds, or the patch bytes change.

2. Seal fixture inputs before the test process starts

Hash only inputs that can change a verdict. Include fixture bytes, the seed, the clock mode, and the locale. Exclude log timestamps, hostnames, and wall-clock duration. Those values vary without changing the contract under test.

sha256sum fixtures/orders.json fixtures/clock.json > review/runner/inputs.sha256
printf 'seed=%s\nclock=%s\nlocale=%s\n' "$SEED" frozen C.UTF-8 >> review/runner/inputs.sha256
sha256sum review/runner/inputs.sha256 | awk '{print $1}' > review/runner/digest.txt
printf 'sealed\n' > review/runner/digest.status
test -s review/runner/digest.txt
Enter fullscreen mode Exit fullscreen mode
import hashlib
from pathlib import Path

def seal_fixture(paths: list[Path], seed: str, clock: str, locale: str) -> str:
    hasher = hashlib.sha256()
    for path in sorted(paths, key=lambda item: item.as_posix()):
        hasher.update(path.as_posix().encode())
        hasher.update(path.read_bytes())
    hasher.update(seed.encode())
    hasher.update(clock.encode())
    hasher.update(locale.encode())
    digest = hasher.hexdigest()
    out = Path("review/runner")
    out.mkdir(parents=True, exist_ok=True)
    (out / "digest.txt").write_text(digest + "\n")
    (out / "digest.status").write_text("sealed\n")
    return digest
Enter fullscreen mode Exit fullscreen mode

Call seal_fixture in the runner job, then start tests in that same job. If a fixture file changes after digest.status is sealed, delete the credit file and the logs, and seal again. Do not attach the old digest to the new bytes.

An open status is not a partial pass. Property results recorded before sealed are discarded, even if they look green.

3. Hide failure logs until that status file exists

This is a visibility rule, not an immutability rule for the whole evidence tree. The runner may still rewrite logs on a fresh job. The author must not read them early.

from pathlib import Path

def logs_visible(status_path: Path) -> bool:
    if not status_path.is_file():
        return False
    return status_path.read_text().strip() == "sealed"
Enter fullscreen mode Exit fullscreen mode

Enforce the barrier at the storage boundary you actually have. A local prototype can use directory permissions. A CI prototype can upload junit.xml only after the status object exists. If the runner cannot delay publication, do not pretend the barrier held.

Treat that run as E_LOGS_EARLY and refuse a freeze. The practical effect is small and strict. The model can propose properties from the diff and the public fixtures. It cannot propose them from the stack trace of the current run.

After the digest is sealed, logs may be published. Any follow-up edit needs a new proposal id. That new id starts the table over.

4. Let the runner score property credit

Credit is a set of changed production symbols that appear in at least one non-tautological property. That property must have passed on the sealed digest. The author job does not get to mark a symbol covered.

A tautology is a case whose assertion cannot fail. A comparison of a value to itself, or a check against a constant the test just wrote, is a typical form. Give those cases zero credit. Classification belongs in the runner, beside execution, so a suggestion cannot bless itself.

def property_credit(credit_id: str, changed: set[str], cases: list[dict], digest: str) -> dict:
    credited = set()
    test_ids = []
    tautologies = []
    for case in cases:
        if case["digest"] != digest or case["status"] != "passed":
            continue
        if case["kind"] == "tautology":
            tautologies.append(case["id"])
            continue
        overlap = set(case["symbols"]) & changed
        if overlap:
            credited |= overlap
            test_ids.append(case["id"])
    missing = sorted(changed - credited)
    return {
        "credit_id": credit_id,
        "digest": digest,
        "credited": sorted(credited),
        "test_ids": test_ids,
        "missing": missing,
        "tautologies": tautologies,
        "pass": bool(changed) and not missing,
    }
Enter fullscreen mode Exit fullscreen mode

Suggested case file, labeled as an example input rather than a captured run:

{
  "changed": ["apply_discount", "tax_cents"],
  "cases": [
    {
      "id": "prop-discount-bounds",
      "digest": "abc123",
      "kind": "property",
      "status": "passed",
      "symbols": ["apply_discount"]
    },
    {
      "id": "prop-tax-identity",
      "digest": "abc123",
      "kind": "tautology",
      "status": "passed",
      "symbols": ["tax_cents"]
    }
  ]
}
Enter fullscreen mode Exit fullscreen mode

On digest abc123, apply_discount is credited. tax_cents is not, because the only matching case is a tautology. pass is false. Stop here.

Filing a freeze does not repair missing credit. One passing property per changed symbol is the minimum in this specification. It is not a coverage target.

Teams with a stronger oracle should require more cases. This article does not set that number. The runner, not the proposal, assigns credit_id when it writes the file.

5. Bind any flake freeze to the credit file

A freeze is optional. Omit it when the required tests passed. When a noisy test remains, the waiver is valid only if every field below matches.

  • runner_digest equals the sealed digest.
  • credit_id points at a credit file with pass true.
  • author_job differs from runner_job.
  • test_id is not one of credit["test_ids"].
  • expires_at is later than the review clock.
from datetime import datetime

def freeze_ok(freeze: dict, credit: dict, now: datetime) -> tuple[bool, str]:
    if freeze["author_job"] == freeze["runner_job"]:
        return False, "E_SHARED_JOB"
    if not credit["pass"]:
        return False, "E_CREDIT_MISSING"
    if freeze["runner_digest"] != credit["digest"]:
        return False, "E_FREEZE_UNBOUND"
    if freeze.get("credit_id") != credit.get("credit_id"):
        return False, "E_FREEZE_UNBOUND"
    if freeze["test_id"] in set(credit.get("test_ids", [])):
        return False, "E_FREEZE_COVERS_CREDIT"
    expiry = datetime.fromisoformat(freeze["expires_at"])
    if expiry <= now:
        return False, "E_FREEZE_EXPIRED"
    return True, "ALLOW"
Enter fullscreen mode Exit fullscreen mode

The sample expiry below is an illustration inside a proposed file. It is not a recommended duration.

{
  "freeze_id": "fz-014",
  "author_job": "reviewer-job-2",
  "runner_job": "server-job-88",
  "runner_digest": "abc123",
  "credit_id": "credit-014",
  "test_id": "test_report_timezone_flake",
  "expires_at": "2026-10-19T00:00:00+00:00"
}
Enter fullscreen mode Exit fullscreen mode

After expiry, delete the waiver. Replay needs a new runner job and a new digest. Matching old fields is not enough if the status file was never sealed for the new bytes.

Stable reject codes

Return codes, not sentences, so CI can branch. Map unknown codes to reject.

Code Condition Required next step
E_SHARED_JOB Author and runner ids match New runner job
E_JOB_MISSING Either id is empty Fill both ids
E_DIGEST_OPEN Status is not sealed Seal, then rerun
E_LOGS_EARLY Logs exist without a sealed status Discard logs
E_CREDIT_MISSING A changed symbol lacks a real property Add a property and reseal
E_FREEZE_UNBOUND Digest or credit id mismatch Drop the freeze
E_FREEZE_COVERS_CREDIT Waiver names a credit-bearing test Narrow the waiver
E_FREEZE_EXPIRED Expiry is not in the future New digest, new waiver
def decide(record: dict, credit: dict, freeze: dict | None, now: datetime) -> str:
    shared = reject_shared_job(record)
    if shared:
        return shared
    if record.get("digest_status") != "sealed":
        return "E_DIGEST_OPEN"
    if record.get("logs_present") and not record.get("logs_after_seal"):
        return "E_LOGS_EARLY"
    if not credit.get("pass"):
        return "E_CREDIT_MISSING"
    if freeze is None:
        return "ALLOW"
    _ok, code = freeze_ok(freeze, credit, now)
    return code
Enter fullscreen mode Exit fullscreen mode
python gate.py --record review/proposal.json --credit review/runner/credit.json --freeze review/freeze.json --now 2026-10-12T12:00:00+00:00
Enter fullscreen mode Exit fullscreen mode

Expect a single code on stdout. ALLOW is the only success value. A checker that fails open will accept a truncated credit file.

Specification trace, not a measured run

Walk the example inputs through the functions above. This trace checks the rule. It is not a benchmark, and it is not a result from a shared service.

The diff changes two production names, apply_discount and tax_cents. Two cases share digest abc123. One is a property and passes. One is marked tautology and passes.

property_credit returns missing: ["tax_cents"] and pass: false. freeze_ok then returns E_CREDIT_MISSING even if the freeze file is otherwise well formed.

Change the tautology into a real bound on tax_cents and keep the same digest. Credit can then pass. If the freeze names prop-discount-bounds, and that id is in credit["test_ids"], the result is E_FREEZE_COVERS_CREDIT.

Point the waiver at test_report_timezone_flake instead, with a future expiry, and the result can be ALLOW. Those branches are the acceptance tests for the checker itself. Run them before you point the checker at agent output.

A gate that cannot reject its own sample inputs is not ready for patches. Store the expected codes next to gate.py, and fail the build if they drift.

Where a free model and a free server fit

Disclosure: This article was prepared as part of MonkeyCode's product outreach.

MonkeyCode's free model access fits the author job: draft the diff, list changed symbols, and suggest property cases from fixtures the runner will hash. MonkeyCode's free server option fits the runner job: execute seal_fixture, run the suite, and write digest.txt plus credit.json. Both are availability options supplied for this workflow. This article does not name models, quotas, hardware, time limits, or scores, because those details were not provided.

Use separate job ids even when both sides are free. reject_shared_job does not care about price. If the free server cannot store digest.status before it publishes logs, keep the run local or pick another runner. Do not file fz-* against an early log.

A restrained way to use that free tier is rehearsal. Ask the free model for YAML that matches the records in this article, review every field, and only then spend a free server run on a sealed fixture set you already trust.

Limitations

A sealed digest makes a run reproducible relative to the hashed inputs. It does not make those inputs correct. A wrong price table, hashed cleanly, still yields a wrong credit decision.

Property credit ignores untested branches inside a credited symbol. The minimum in section 4 can admit a thin test. Raise it in your own policy if the diff is wide, and do not import a universal count from this text.

Tautology detection here is a label, kind == "tautology". A real runner needs a classifier. If that classifier misses an equality where both sides come from the patch, credit will be inflated. Review suggested properties before they enter the sealed set.

Freezes create debt. The date in the sample JSON is only a placeholder so the parser has a value to read. Choose an expiry your rotation will actually revisit. An unreviewed renewal is the same loop this gate is meant to stop.

Free server execution is a trust boundary. Confirm who can edit fixtures, how long logs are kept, and whether other tenants can read them before you upload proprietary tests. This article does not verify those controls.

Symbol extraction is the sharp edge. A parser that misses renames will under-count credit and block safe patches. A parser that reads comments will over-count credit and admit weak patches. Calibrate that parser on your repository, not on the two-function example above.

Who should skip this workflow

Do not adopt the gate if you cannot issue two job ids and delay log publication. The interesting checks all assume that split.

Skip it when tests stay non-deterministic after seed, clock, and fixture bytes are locked. A freeze would record the noise and give it an expiry. Fix the source first.

Skip it during an incident that cannot wait for a reseal. Use the incident path, then backfill a digest after recovery. Do not mint a retroactive freeze for a log that was readable before sealing.

Skip it for changes the runner cannot see, such as configuration applied only in production. A green credit file on the wrong tree is a sealed irrelevance.

Skip it if you need a regulated validation package or a formal proof. The checker is a review aid. It is not a certification.

What the reviewer should be able to answer

Four questions should be answerable from files, without asking the model.

  1. Which job wrote the patch?
  2. Which job wrote digest.status?
  3. Which changed symbols have non-tautological credit on that digest?
  4. If a freeze file exists, does it cite that digest, a passing credit id, a different test, and a future expiry?

A missing answer keeps the patch proposed. A free drafting slot and a free execution slot do not retire the questions. They only decide where the proposal is cheap to write, and where the sealed run is cheap to repeat.

Top comments (0)