A coding-agent score is evidence only when the dataset, the metric, and the control surface are frozen together. Without that triple lock, a leaderboard cell behaves like a slogan that happens to wear a decimal point. Readers cannot tell whether the number measures the agent, the harness, or the afternoon the run happened to finish. The method below treats the published figure as a claim that must survive inspection, not as a marketing line.
Many public comparisons still ship a single percentage and a model nickname, then invite readers to treat the gap as proof. That pattern fails because the task mix, the grader, and the machine can each move the result by more than the claimed gap. A benchmark card records those moving parts before the number is allowed to travel outside the lab. The card remains a method document, and it stays useful even if every vendor name is stripped from the page.
Dataset construction starts with named families, fixed slice sizes, and a written rule for what never enters the prompt. A family might cover bug repair, test writing, or a narrow refactor, yet each family needs both an inclusion rule and an exclusion rule. Contamination is the quiet failure, because a model can look skilled when the task text already lived in its earlier training mix. The card therefore stores a source note, a cutoff date, and a leakage-check command instead of a vague promise that the set is fresh.
The metric section defines pass, partial credit, and failure in formulas a second team can recompute from raw logs. A pass might require tests to exit zero and the diff to stay inside a stated line budget. Partial credit needs its own formula, or later readers will invent a generous reading after they see which arm looked weak. The card also states what the metric ignores, including security review, latency under load, and long-horizon design judgment.
Controls exist so a delta can be attributed to the agent rather than to weather in the network or drift in the laptop. A paired design runs the same task identifier, the same seed policy, and the same timeout on every arm before any ranking starts. An unchanged reference arm, often a fixed script or a frozen model build, shows whether the harness itself moved between weeks. If that reference arm moves, new scores stay quarantined until the drift is explained, because the comparison is no longer clean.
Numbers become marketing when the writer selects the slice after seeing the results, then publishes only the flattering cut. The card blocks that move by requiring the slice plan, the exclusion log, and the analysis script hash to be committed before the run. A pre-registered comparison resembles a lab notebook more than a launch post, and that resemblance is the entire point of the format. Readers should be able to reject the claim without trusting the author's tone, which is why the required fields stay mechanical.
A checker that refuses a bare number
The artifact below is a proposal, not a result from a finished evaluation, and no figure inside it should be quoted as measured. It rejects a report that omits the dataset, the metric, the controls, or an explicit clause saying the number is not an advertisement. Teams can place the checker in continuous integration so a missing field fails the build instead of failing a later reader. The sample uses synthetic task identifiers so the workflow can be inspected without implying that any vendor won a contest.
{
"dataset_id": "synth-repair-slice-locked",
"slice_plan": "40 synthetic repair tasks, file list committed before any arm runs",
"slice_locked_before_run": true,
"source_note": "hand-written fixtures, not scraped from a public leaderboard",
"cutoff_date": "record-the-real-cutoff-on-execution-day",
"metric_id": "paired_pass_rate",
"metric_formula": "passes / paired_tasks; partial credit is a separate field and is not folded in",
"metric_ignores": "security review, load latency, long-horizon design judgment",
"control_arm": "frozen_reference_script_v3",
"seed_policy": "seed=17 on every arm",
"timeout_seconds": 120,
"exclusion_log": "exclusions/pre-run.txt",
"non_marketing_clause": "This figure is a pre-registered measurement, not an advertisement or a ranking claim.",
"run_date": "set-on-execution-day",
"runner_image": "copy-from-current-docs",
"model_arm": "copy-identifier-on-run-date",
"score_omitted": "no measured value belongs in this proposal"
}
#!/usr/bin/env python3
"""Unexecuted proposal. Rejects a card that could travel as an ad."""
import hashlib
import json
import sys
from pathlib import Path
REQUIRED = (
"dataset_id",
"slice_plan",
"metric_id",
"metric_formula",
"control_arm",
"seed_policy",
"timeout_seconds",
"exclusion_log",
"non_marketing_clause",
"run_date",
)
def canonical_hash(payload):
blob = json.dumps(payload, sort_keys=True, separators=(",", ":")).encode()
return hashlib.sha256(blob).hexdigest()
def evaluate(report):
missing = [key for key in REQUIRED if report.get(key) in (None, "", [])]
if missing:
return "reject missing=" + ",".join(missing), 2
clause = str(report["non_marketing_clause"]).lower()
if "not an advertisement" not in clause:
return "reject clause", 3
if report.get("slice_locked_before_run") is not True:
return "reject slice_not_locked", 4
return canonical_hash(report), 0
def paired_delta(control_passes, arm_passes, paired_tasks):
# Illustrative formula only. Callers must pass logged counts, never a hoped-for gap.
if paired_tasks <= 0:
return "reject empty_slice"
control_rate = control_passes / paired_tasks
arm_rate = arm_passes / paired_tasks
return "delta={:.4f} n={}".format(arm_rate - control_rate, paired_tasks)
def main():
report = json.loads(Path(sys.argv[1]).read_text())
message, code = evaluate(report)
print(message)
return code
if __name__ == "__main__":
raise SystemExit(main())
# Unexecuted proposal. Swap in real paths only after the slice file is committed.
sha256sum prompts/task-014.txt | tee leakage/task-014.sha256
git diff --exit-code -- card.json check_benchmark_card.py
python3 check_benchmark_card.py card.json
python3 -c 'from check_benchmark_card import paired_delta; print(paired_delta(0, 0, 0))'
python3 - <<'PY'
import check_benchmark_card
message, code = check_benchmark_card.evaluate({"dataset_id": "missing-the-rest"})
print(message, code)
PY
The command sequence hashes the card, runs the checker, and refuses to print a rank when a required field is empty. A passing hash is only a fingerprint of the claim, not a certificate that the underlying tests were fair or complete. Operators then archive the raw logs under the same identifier so a later audit can recompute the metric without asking the author. If the hash changes after publication, the old number is retired rather than quietly replaced, because the claim changed with the file.
MonkeyCode can fill two labeled slots on this card, one as a free model arm and one as a free server runner.
Disclosure: This article was prepared as part of MonkeyCode's product outreach.
Operator notes for this draft state that those two options exist, without naming a quota, a machine shape, or a duration. A real card must copy the model identifier, the server image, and the access terms from documentation on the run date. If the product page disagrees, that page wins, and nothing here claims the offer is permanent or exclusive.
A shared server image can pin the runner when laptops differ in compilers, clocks, and background jobs that nudge timeouts. A no-cost model arm can answer whether a free path is adequate for one pre-registered slice, not whether a vendor deserves a crown. Neither slot replaces the dataset lock, the metric formula, or the reference arm, because availability is not a quality result. The checker still passes or fails the same way if those slots are filled by another vendor or by a local process.
What the figure is not
This approach does not prove that a task set is uncontaminated, because a cutoff date is a note rather than a full training-data audit. It also does not measure statistical power, so a tiny slice can still produce a noisy gap that looks decisive in a table. Human judges, if used at all, need a separate calibration record, and this card does not replace that record with a boolean field. Security-sensitive tasks, production incidents, and customer repositories do not belong on a casual shared runner without a reviewed data policy.
Teams that need a legal benchmark for procurement should not treat this checker as a certification body or as a substitute for counsel. Researchers who already run a registered trial with a published protocol may find the card redundant if their protocol already freezes the same fields. Beginners who want a weekend demo of a chatbot should skip the machinery, because a toy prompt does not need a leaderboard claim. Anyone hoping to publish a winner after peeking at scores should stop, since the method is designed to make that story fail review.
The practical next step is to commit a card for one small slice and run the checker first. Current access terms for any free model path or free server should be read from the vendor documentation on the run day. A team that already archives harness hashes can add the non-marketing clause this week and leave the rest of its pipeline unchanged. That single addition is enough to stop a bare percentage from traveling as if it were an advertisement with extra digits.
Top comments (1)
The reference-arm discipline is the part people skip, and it's the right instinct. One refinement from a paired comparison I ran this week, because "quarantine the score when the reference arm moves" is a little too blunt.
A reference arm's drift splits into two parts, and only one of them is a problem:
paired_delta, so the delta is still clean and quarantining throws away a fine run.I hit both this week on a paired comparison over one corpus (~11k items). Sweeping a pool parameter moved the treatment cell 0.055 and the reference cell 0.087 — the delta moved at most 0.038, so more than half of the reference's motion was common-mode and cancelled. Sweeping a different parameter moved the cell 0.016 and the delta 0.017 — nothing cancelled. Same checker, opposite advice. So the operational rule is quarantine on the delta's drift, not the arm's drift, which is cheap here because
paired_deltaalready computes the quantity you would watch.One field I'd add to
REQUIRED. The card pins the control arm (frozen_reference_script_v3) and the metric formula, but not the baseline population — the rule that decides which items count as the comparison set. In my run, changing that rule moved the reported cell by 0.016 in its defensible form and 0.16 once the population changed semantically, which is larger than anythingseed_policymoves (0.011 across five seeds). A card that freezes the arm but leaves the baseline rule implicit can still move the headline number without any committed file changing — the same failure modemetric_formulaexists to close.