An agent score becomes evidence only after the trial budget, the difficulty strata, and the exclusion rules are frozen. A single clean transcript can still guide a repair, yet one lucky path does not estimate error. The method below treats the headline percentage as a summary of paired trials, not as a brochure claim. A reviewer can reproduce the accounting with a small Python harness before any model name enters the discussion.
A clean run is not a timetable
A clean local run often looks decisive, especially when every visible check happens to pass on the first afternoon. That result can guide the next edit, yet it remains a weather report until the same items repeat under fixed rules. A transit agency does not publish a timetable from one sunny commute, and a lab should not publish a rank from one polished run. The analogy is deliberate, because variance, route difficulty, and cancelled trips all belong in the published record.
What the pack must freeze
The dataset is a frozen task pack split into difficulty strata before any candidate is allowed to run. Each stratum holds tasks of similar expected effort, such as a one-file fix, a multi-file refactor, or a failing integration test. Items carry stable identifiers, and the pack records the stratum label beside the prompt path and the oracle command. A later chart may omit a stratum, but the scorer may not silently move an item into an easier bin.
Without strata, a model that excels at short edits can outscore a model that survives harder repairs, and the average hides the swap. Promotional copy prefers that average, because a single percentage travels farther than a table of uneven wins. The methodology therefore reports a within-stratum paired difference, and only then a predeclared summary across strata. If the summary and the strata disagree, the strata remain the result and the summary becomes a footnote.
Metrics that refuse a slogan
The primary metric is the paired outcome on the same item, not an unpaired accuracy computed on different samples. Each trial records pass or fail against the oracle, the wall time, and whether the harness aborted for a logged reason. A secondary metric is the stratum win rate, defined as the share of items where the candidate beats the control. Neither metric is a quality slogan, because both inherit every exclusion that the log has already accepted.
Controls exist to stop the number from drifting into advertising copy after the run looks favorable. The control arm uses the same harness, the same oracle, and the same retry ceiling as the candidate arm. A placebo item with an impossible oracle must fail, or the scorer treats the whole batch as untrustworthy. That negative control is the fire alarm, not an optional appendix for readers who enjoy methodology notes.
The trial budget is chosen before scores appear, and it states how many independent repeats each item receives. Two repeats can reveal a flip, while a larger budget is required before a small gap deserves a public rank. This draft does not prescribe a universal sample size, because task variance differs across repositories and oracles. It does require the chosen budget to be written into the manifest, so a disappointing repeat cannot be deleted quietly.
Exclusions are the place where an honest lab and a promotional writeup usually diverge in practice. A pre-registered rule may drop an item for a missing fixture, a leaked solution, or an oracle that no longer builds. A post-hoc rule that drops an item because the candidate failed is not an exclusion, and the harness should refuse it. The log stores the rule identifier, the item identifier, and the timestamp, so a reviewer can replay the decision.
A harness that enforces the freeze
The following Python module is a proposed harness, and it has not been executed against a live scoring service. It hashes the manifest, checks paired outcomes, and rejects any exclusion that lacks a pre-registered rule. Operators should treat the script as a starting ledger, then replace the toy fields with their own oracle records. The point is the accounting, not a claim that any particular model already passed these checks.
#!/usr/bin/env python3
"""Proposed trial-budget harness.
This example has not been executed against a live scoring service.
Replace the fixture files before using it on a real task pack.
"""
from __future__ import annotations
import hashlib
import json
import sys
from pathlib import Path
RULES = {
"missing_fixture": "Oracle inputs were absent before the candidate started.",
"leaked_solution": "A reference patch was visible in the prompt context.",
"oracle_broken": "The oracle command failed on the untouched baseline.",
}
def manifest_hash(path: Path) -> str:
return hashlib.sha256(path.read_bytes()).hexdigest()
def load_manifest(path: Path) -> dict:
data = json.loads(path.read_text(encoding="utf-8"))
required = {"pack_id", "trial_budget", "strata", "items", "exclusion_rules"}
missing = sorted(required.difference(data))
if missing:
raise SystemExit(f"manifest missing fields: {missing}")
budget = data["trial_budget"]
if not isinstance(budget, int) or budget < 1:
raise SystemExit("trial_budget must be a frozen integer of at least 1")
return data
def accept_exclusion(rule_id: str, registered: list[str]) -> None:
if rule_id not in registered or rule_id not in RULES:
raise SystemExit(f"post-hoc exclusion refused: {rule_id}")
def paired_delta(candidate: list[bool], control: list[bool]) -> dict:
if len(candidate) != len(control) or not candidate:
raise SystemExit("paired arms require equal non-empty outcomes")
wins = losses = ties = 0
for cand, base in zip(candidate, control):
if cand == base:
ties += 1
elif cand and not base:
wins += 1
else:
losses += 1
return {"wins": wins, "losses": losses, "ties": ties, "n": len(candidate)}
def self_check() -> None:
delta = paired_delta([True, False], [False, False])
expected = {"wins": 1, "losses": 0, "ties": 1, "n": 2}
if delta != expected:
raise SystemExit("paired_delta self-check failed")
try:
accept_exclusion("candidate_failed", ["missing_fixture"])
except SystemExit:
print("self-check ok")
return
raise SystemExit("self-check failed to refuse a post-hoc exclusion")
def evaluate(manifest_path: Path, outcomes_path: Path) -> dict:
manifest = load_manifest(manifest_path)
digest = manifest_hash(manifest_path)
outcomes = json.loads(outcomes_path.read_text(encoding="utf-8"))
if outcomes.get("manifest_sha256") != digest:
raise SystemExit("outcome file does not match frozen manifest hash")
for event in outcomes.get("exclusions", []):
accept_exclusion(event["rule_id"], manifest["exclusion_rules"])
by_stratum: dict[str, list[dict]] = {}
for item in manifest["items"]:
arm = outcomes["items"][item["id"]]
if len(arm["candidate"]) != manifest["trial_budget"]:
raise SystemExit(f"budget mismatch: {item['id']}")
if len(arm["control"]) != manifest["trial_budget"]:
raise SystemExit(f"control budget mismatch: {item['id']}")
if item.get("placebo") and any(arm["candidate"]):
raise SystemExit(f"placebo passed: {item['id']}")
by_stratum.setdefault(item["stratum"], []).append(
paired_delta(arm["candidate"], arm["control"])
)
return {
"manifest_sha256": digest,
"pack_id": manifest["pack_id"],
"strata": by_stratum,
}
def main() -> None:
if len(sys.argv) == 2 and sys.argv[1] == "self-check":
self_check()
return
if len(sys.argv) != 3:
raise SystemExit(
"usage: trial_budget.py self-check | trial_budget.py PACK OUTCOMES"
)
report = evaluate(Path(sys.argv[1]), Path(sys.argv[2]))
print(json.dumps(report, indent=2))
if __name__ == "__main__":
main()
A reviewer can freeze the pack with a checksum before the first candidate process is allowed to start. The commands below assume a local checkout, a Python 3 interpreter, and no network call inside the toy oracle. They illustrate the control path, and they do not measure a hosted product or imply a hardware specification. If a command fails, the failure stays in the log rather than disappearing into a revised average.
python3 trial_budget.py self-check
python3 -c "import hashlib, pathlib; print(hashlib.sha256(pathlib.Path('task_pack.json').read_bytes()).hexdigest())"
python3 trial_budget.py task_pack.json outcomes.json
The sample manifest freezes a two-stratum pack, a trial budget of two, and the three exclusion rules the harness recognizes. The sample outcomes file must carry the same manifest hash, or the script exits before it prints a table. A placebo item that passes on the candidate arm aborts the batch, which is the intended failure for a broken scorer. Those files are fixtures for the method, and they are not measurements of any hosted coding agent.
{
"pack_id": "strata-demo-001",
"trial_budget": 2,
"strata": ["short_edit", "multi_file"],
"exclusion_rules": ["missing_fixture", "leaked_solution", "oracle_broken"],
"items": [
{
"id": "short-01",
"stratum": "short_edit",
"prompt_path": "prompts/short-01.md",
"oracle": "pytest -q tests/test_short_01.py",
"placebo": false
},
{
"id": "placebo-oracle",
"stratum": "short_edit",
"prompt_path": "prompts/placebo.md",
"oracle": "false",
"placebo": true
}
]
}
{
"manifest_sha256": "<paste sha256 of task_pack.json>",
"exclusions": [],
"items": {
"short-01": {"candidate": [true, false], "control": [false, false]},
"placebo-oracle": {"candidate": [false, false], "control": [false, false]}
}
}
Where a hosted option fits
Disclosure: This article was prepared as part of MonkeyCode's product outreach.
MonkeyCode's free model access and free server option matter only as the candidate runtime, while the ledger stays local. That availability is only an operator-supplied convenience, and this draft does not verify quotas, hardware, duration, or model names. A team may copy the limits it actually sees into the run notes, without treating free capacity as a quality score.
Why the percentage is not marketing
A percentage becomes marketing when the audience cannot see the pack, the budget, the exclusions, or the negative control. The same percentage becomes a lab result when those four artifacts travel with the figure and a reviewer can recompute it. Publishing the strata beside the summary also blocks the common slide that hides a loss on hard items. If a stakeholder wants only the largest number, the writeup should withhold the rank rather than inflate the claim.
Limits, and who should walk away
This approach does not detect training-set contamination beyond the canary items the pack author remembered to include. It does not price the runs, and a cheaper failure can still be the better engineering choice for a given team. Wall-clock comparisons remain fragile when a free runner and a paid runner differ in queueing, which this toy script cannot see. Small strata will swing wildly, so a two-item bin should be reported as a case series rather than as a rate.
Teams that need a vendor ranking for a launch post this week should not use this harness as decoration. The method also fits poorly when the oracle itself is disputed, because a shared wrong test will crown the wrong patch. Researchers who cannot freeze the task pack, or who must tune prompts after seeing scores, will violate the exclusion rule immediately. In those settings the honest output is a demo transcript with a date, not a benchmark table.
The practical next step is to hash a small pack, run the placebo item, and confirm that a bad exclusion is rejected. Only after that check should a candidate arm consume either a local process or a hosted free option. The resulting table will look quieter than a leaderboard, which is the intended trade for a number that can survive review. Quiet evidence is still the result a later reader can trust when the demo transcript has already scrolled out of view.
Top comments (0)