A published coding-agent score describes one frozen task slice, and it does not describe an entire product. That slice needs a named dataset, an explicit metric contract, and controls a later reader can rerun. Without those three artifacts beside the number, the percentage behaves like a poster rather than evidence. Careful readers then import a marketing reading that the measurement itself never earned and cannot support.
Recent public notes on agent rankings often print one resolved rate and then stop the explanation entirely. The missing piece is not a longer adjective list, but the identity of the tasks behind that rate. A rate computed on easy stubs will flatter a system that fails on multi-file repairs with hidden tests. That gap is why a benchmark card belongs in the same commit as the score, rather than in a later aside.
What the card records
The dataset section names the source repository, the pinned commit, and the inclusion rule used during sampling. Tasks enter the slice only when they carry a failing test, a bounded edit surface, and a license that permits prompt redistribution. Stratification by language and by estimated edit span keeps a Python-heavy week from masquerading as a general coding result. Each retained task also stores the commit that produced the failing test, so a later rerun can detect an upstream rewrite.
The metric contract defines the numerator, the denominator, and the attempt budget before any model is called. Resolved rate counts a task only when hidden tests pass within that budget, or when a passing patch arrives earlier. Pass-at-k stays in a separate column, because a lucky later attempt is not the same event as a first repair. Duration and token totals can sit beside quality, yet they do not replace the contract that says what resolved means.
Controls exist so the number can fail, which is the property that separates a benchmark from a brochure. A no-edit control submits the untouched repository and must stay near zero on tasks that truly require a patch. A known-fix control applies a recorded one-line repair and must pass, proving the harness actually executes the hidden tests. If either control misbehaves, the run is discarded rather than averaged into a leaderboard that nobody can later audit.
Marketing copy selects the friendliest slice after the fact and then narrates the winner as if that slice had been fixed. A benchmark card reverses the order by recording the slice first and by refusing publication when the recorded slice no longer matches. The match check does not prove that a model is strong, and it only proves the claim still points at the same tasks. Readers can then dispute the metric without guessing which repositories were quietly dropped from the denominator.
Freeze the slice locally
The module below is a proposal for that card, and it has not been executed against a live agent fleet. It writes a JSON manifest, checks stratum floors, and exits non-zero when the slice no longer matches the recorded digest. Operators may rename fields, but they should not delete the refusal path that blocks a drifted publish. The example uses a local JSON file so the method can be reviewed without a network call or a vendor SDK.
#!/usr/bin/env python3
"""Proposal: freeze a benchmark card. Not executed against a live agent fleet."""
from __future__ import annotations
import hashlib
import json
import sys
from pathlib import Path
REQUIRED = (
"task_id",
"repo",
"commit",
"language",
"edit_span",
"license",
"has_failing_test",
)
ALLOWED_LICENSES = {"mit", "apache-2.0", "bsd-3-clause"}
MIN_PER_STRATUM = 5
ATTEMPT_BUDGET = 1
def load_tasks(path: Path) -> list[dict]:
tasks = json.loads(path.read_text(encoding="utf-8"))
if not isinstance(tasks, list) or not tasks:
raise SystemExit("task file must be a non-empty list")
for task in tasks:
missing = [field for field in REQUIRED if field not in task]
if missing:
raise SystemExit(f"{task.get('task_id', '?')} missing {missing}")
if task["has_failing_test"] is not True:
raise SystemExit(f"{task['task_id']} lacks a failing test")
if task["license"] not in ALLOWED_LICENSES:
raise SystemExit(f"{task['task_id']} license is outside the inclusion rule")
if not isinstance(task["edit_span"], int) or task["edit_span"] < 1:
raise SystemExit(f"{task['task_id']} has a non-positive edit span")
return tasks
def stratum_key(task: dict) -> str:
bucket = "narrow" if task["edit_span"] <= 30 else "wide"
return f"{task['language']}:{bucket}"
def stratum_counts(tasks: list[dict]) -> dict[str, int]:
counts: dict[str, int] = {}
for task in tasks:
key = stratum_key(task)
counts[key] = counts.get(key, 0) + 1
short = {key: count for key, count in counts.items() if count < MIN_PER_STRATUM}
if short:
raise SystemExit(f"stratum floor missed: {short}")
return counts
def digest_of(tasks: list[dict]) -> str:
blob = json.dumps(tasks, sort_keys=True, separators=(",", ":")).encode("utf-8")
return hashlib.sha256(blob).hexdigest()
def build_card(tasks: list[dict]) -> dict:
return {
"dataset_digest": digest_of(tasks),
"task_count": len(tasks),
"strata": stratum_counts(tasks),
"metric": {
"name": "resolved_rate",
"numerator": "hidden tests pass within the attempt budget",
"denominator": "tasks in this frozen slice",
"attempt_budget": ATTEMPT_BUDGET,
"pass_at_k_reported_separately": True,
},
"controls": ["no_edit_must_fail", "known_fix_must_pass"],
"not_a_claim_about": [
"unseen repositories",
"security behavior",
"runner permanence",
],
}
def freeze(tasks: list[dict], dest: Path) -> None:
card = build_card(tasks)
dest.write_text(json.dumps(card, indent=2) + "\n", encoding="utf-8")
print(f"froze {card['dataset_digest'][:12]} tasks={card['task_count']}")
def verify(tasks: list[dict], card_path: Path) -> None:
recorded = json.loads(card_path.read_text(encoding="utf-8"))
actual = digest_of(tasks)
if recorded.get("dataset_digest") != actual:
raise SystemExit("slice drifted; refuse to publish the score")
stratum_counts(tasks)
print(f"slice ok {actual[:12]} tasks={len(tasks)}")
def main() -> None:
if len(sys.argv) != 4:
raise SystemExit("usage: benchmark_card.py freeze|verify tasks.json card.json")
command, task_path, card_path = sys.argv[1], Path(sys.argv[2]), Path(sys.argv[3])
tasks = load_tasks(task_path)
if command == "freeze":
freeze(tasks, card_path)
elif command == "verify":
verify(tasks, card_path)
else:
raise SystemExit("use freeze or verify")
if __name__ == "__main__":
main()
A minimal task file can live next to the module, with one object per repository repair the harness is allowed to score. Each object records language, edit span, license, and whether a failing test already exists before the agent runs. The stratum floor of five and the attempt budget of one are example thresholds in this proposal, not universal standards. A one-task file should fail freeze, and that failure demonstrates the floor instead of a broken script.
[
{
"task_id": "demo-1",
"repo": "example/widgets",
"commit": "abc123",
"language": "python",
"edit_span": 12,
"license": "mit",
"has_failing_test": true
}
]
python3 benchmark_card.py freeze tasks.min.json card.json
python3 benchmark_card.py verify tasks.min.json card.json
The second script is a separate publish gate, and it is labeled unexecuted because no live outcomes were supplied here. It reads a local report, demands both harness controls, and only then computes a resolved rate. The printed rate stays tied to the attempt budget in that report, so a later plot cannot silently change the denominator. Operators should treat a missing control flag as an incomplete run, not as a zero that can be averaged away.
#!/usr/bin/env python3
"""Unexecuted proposal: refuse a score unless both harness controls passed."""
import json
import sys
from pathlib import Path
def resolved_rate(outcomes: list[dict], budget: int) -> float:
eligible = [row for row in outcomes if row["attempts"] <= budget]
if not eligible:
raise SystemExit("no eligible attempts")
passed = sum(1 for row in eligible if row["hidden_tests_passed"])
return passed / len(eligible)
def main() -> None:
report = json.loads(Path(sys.argv[1]).read_text(encoding="utf-8"))
if report.get("no_edit_must_fail") is not True:
raise SystemExit("no-edit control failed; do not publish")
if report.get("known_fix_must_pass") is not True:
raise SystemExit("known-fix control failed; do not publish")
rate = resolved_rate(report["outcomes"], report["attempt_budget"])
print(f"resolved_rate={rate:.3f} n={len(report['outcomes'])}")
print("not a product verdict; read not_a_claim_about in the card")
if __name__ == "__main__":
main()
python3 publish_gate.py report.json
Availability does not create the claim
MonkeyCode's free model access and free server option can host the harness while a team is still arguing about the metric contract. Disclosure: This article was prepared as part of MonkeyCode's product outreach. The free access matters only as a way to repeat the same frozen slice while the card is still under review. A free server option can keep the harness and the manifest together so the digest check is not a laptop-only ritual.
This draft does not name models, quotas, hardware, duration, or permanence, because those details were not verified here. Current product documentation should be read before a run is scheduled, and observed limits should be copied into the card. Neither availability option selects the dataset, defines the resolved rate, or repairs a broken no-edit control by itself. A score produced on a convenient runner still needs the card, since convenience is not a substitute for a frozen claim.
This protocol does not measure security behavior, long-horizon planning, or skill on private monorepos absent from the slice. Hidden tests can leak through verbose logs, so the harness should redact failure output before any retry prompt is built. Small slices carry wide uncertainty, and a three-point gap on forty tasks is often noise rather than a stable ranking. Human review remains necessary when tests pass by deleting assertions or by weakening the oracle that was supposed to judge the patch.
Teams that need a procurement number by Friday should not wear this card as a costume for a decision already made. Researchers who cannot publish the task list, or a durable pointer to it, cannot invite others to audit the rate. Agents that only chat, and never emit a patch against a repository, sit outside a resolved-rate contract entirely. Operators without a failing-test oracle should choose another metric instead of forcing this harness to invent one.
A useful reading of the finished card sounds like a lab note rather than a product launch sentence. It says which languages were thick enough to interpret, which strata missed the floor and were therefore withheld, and which control failed. It also says what the number is not, including unseen repositories, security work, and any claim that a runner remains available. That negative space is what keeps a later slide from promoting a slice result into a universal ranking.
The practical next step is to freeze at least a few dozen tasks, run both controls, and only then attach a model score. Readers who need a shared runner while they rehearse that order can try MonkeyCode's free model access and free server option. The manifest should still land in version control beside the score, because a convenient runner does not audit itself. Until the card and the controls travel with the percentage, the number remains a headline instead of evidence.
Top comments (0)