A leaderboard screenshot lands in your team channel. One coding agent solves 74% of real issues, says the chart. Nobody in the channel has the task list, the commit SHA, the token budget, or the stop condition behind that number. Two weeks later it is in a slide deck; a quarter later it is the reason you picked a vendor.
The fix is smaller than most people expect: a task list, one harness file, and a ledger you can hand to someone else. This article builds that, then runs it on free infrastructure.
Why a published score is a manufacturing process
A benchmark number is not a lie. It is the output of a process where someone chose the tasks, the budget per task, the retry policy, and the fail condition. Those choices are all legitimate. Together they mean the score answers one narrow question: what does this agent do under conditions we picked?
When you cite it, you are implicitly claiming those conditions match yours. Usually they do not. So before you argue about scores, reproduce one.
Step 1 — Fingerprint the claim before you run anything
You cannot reproduce a number you have not described. Write down six fields. If any field is missing from the source, mark it unknown and treat the score as directional only.
| Field | Why it changes the score |
|---|---|
| Task set + version | Different repos, different difficulty |
| Budget per task | More tokens, more retries, higher pass rate |
| Attempts per task | pass@1 vs pass@5 are different metrics |
| Harness commit | Prompt, tools, and stop logic all move the number |
| Model + endpoint | A hosted alias can change under you |
| Verify command | The definition of "solved" |
Do this on paper first. It takes ten minutes and saves a week.
Step 2 — Build the smallest harness that can be re-run
Two files. Everything that can move the score lives in config, not in code.
{
"model_id": "REPLACE_ME",
"base_url": "REPLACE_ME",
"task_dir": "tasks",
"max_tokens_per_task": 6000,
"max_wall_clock_seconds": 180,
"retries": 0,
"seed": 20261008,
"tasks": ["issue-001", "issue-002"]
}
Now the runner. Note the two design rules in the docstring: the harness records its own hash, and it never decides pass/fail — the task's own verify command does.
#!/usr/bin/env python3
"""reproduce.py - rerun a fixed task list, one ledger row per attempt.
Design rules:
* every knob that can change the score lives in eval.config.json
* the harness records its own hash, so an edited harness is visible
* the harness never decides pass/fail; the task's verify command does
"""
from __future__ import annotations
import hashlib
import json
import os
import pathlib
import platform
import subprocess
import sys
import time
ROOT = pathlib.Path(__file__).resolve().parent
CFG = json.loads((ROOT / "eval.config.json").read_text())
RUN_ID = os.environ.get("RUN_ID") or time.strftime("%Y%m%dT%H%M%SZ", time.gmtime())
LEDGER = ROOT / "runs" / f"{RUN_ID}.jsonl"
def fingerprint() -> dict:
return {
"run_id": RUN_ID,
"harness_sha256": hashlib.sha256(
pathlib.Path(__file__).read_bytes()).hexdigest()[:16],
"python": sys.version.split()[0],
"platform": f"{platform.system()}/{platform.machine()}",
"model_id": CFG["model_id"],
"base_url_host": CFG["base_url"].split("/")[2],
"max_tokens_per_task": CFG["max_tokens_per_task"],
"max_wall_clock_seconds": CFG["max_wall_clock_seconds"],
"retries": CFG["retries"],
"seed": CFG["seed"],
}
def patch_hash(workdir: pathlib.Path):
diff = subprocess.run(["git", "diff", "--binary"], cwd=workdir,
capture_output=True)
return hashlib.sha256(diff.stdout).hexdigest()[:16] if diff.stdout else None
def run_one(task_id: str) -> dict:
workdir = ROOT / CFG["task_dir"] / task_id
spec = json.loads((workdir / "task.json").read_text())
env = {
**os.environ,
"MODEL_ID": CFG["model_id"],
"BASE_URL": CFG["base_url"],
"MAX_TOKENS": str(CFG["max_tokens_per_task"]),
"RETRIES": str(CFG["retries"]),
"SEED": str(CFG["seed"]),
}
started = time.monotonic()
agent = subprocess.run(
[sys.executable, str(ROOT / "agent.py"), "--task", str(workdir)],
capture_output=True, text=True,
timeout=CFG["max_wall_clock_seconds"], env=env,
)
verify = subprocess.run(
spec["verify"], shell=isinstance(spec["verify"], str),
cwd=workdir, capture_output=True, text=True,
timeout=spec.get("verify_timeout", 300),
)
usage = {}
usage_file = workdir / "usage.json" # your agent wrapper writes this
if usage_file.exists():
usage = json.loads(usage_file.read_text())
return {
"task_id": task_id,
"passed": verify.returncode == 0,
"verify_cmd": spec["verify"],
"verify_stderr_tail": verify.stderr[-400:],
"agent_exit": agent.returncode,
"wall_clock_s": round(time.monotonic() - started, 2),
"tokens_in": usage.get("tokens_in"),
"tokens_out": usage.get("tokens_out"),
"patch_sha256": patch_hash(workdir),
}
def main() -> int:
LEDGER.parent.mkdir(exist_ok=True)
with LEDGER.open("a") as fh:
fh.write(json.dumps({"_fingerprint": fingerprint()}) + "\n")
for task_id in CFG["tasks"]:
row = run_one(task_id)
fh.write(json.dumps(row) + "\n")
print(f"{task_id:>12} {'PASS' if row['passed'] else 'FAIL'} "
f"{row['wall_clock_s']:>7.2f}s "
f"in={row['tokens_in']} out={row['tokens_out']}")
print(f"\nledger: {LEDGER.relative_to(ROOT)}")
return 0
if __name__ == "__main__":
raise SystemExit(main())
One field is deliberately fragile: tokens_in / tokens_out come from a usage.json your agent wrapper writes. Providers report usage differently. Do not scrape stdout for token counts — that is how silent zeros enter a ledger.
The harness needs Python 3.10+ for the union type on patch_hash. Drop the annotation if you are on 3.9.
Step 3 — Pick a substrate you do not have to pay for
A reproduction run is bursty: idle most of the week, then 20 tasks in an hour. That is the worst shape for a monthly cloud bill.
This is where MonkeyCode fits the method rather than the marketing. MonkeyCode is an open-source project that offers free model access and a free server option. Both of those availability claims come from the project operator; this article did not test them, and nothing below depends on them being permanent.
Disclosure: This article was prepared as part of MonkeyCode's product outreach.
Because the claims are operator-supplied, treat them as configuration to verify, not facts to assume. Before your first run, check the project's current docs for: which models the free access currently covers, what the free server currently provides, and what limits or expiry apply. This article asserts none of those specifics, on purpose — quotas, hardware, and duration are exactly the fields that change, and an article is a terrible place to freeze them.
Then keep key hygiene boring:
# Never in the repo. Never in the ledger. Never in a screenshot.
export MONKEYCODE_API_KEY="..."
export RUN_ID=$(date -u +%Y%m%dT%H%M%SZ)
python reproduce.py
The fingerprint logs base_url_host, not the key. That is intentional: you want a teammate to see which endpoint produced the run, and you never want the ledger to be a secret store.
Step 4 — Turn the ledger into one honest summary
Raw JSONL is not a score. Reduce it once, in code, so nobody reduces it by hand in a spreadsheet.
#!/usr/bin/env python3
"""score.py - one ledger in, one honest summary out."""
import json
import pathlib
import sys
rows = [json.loads(l) for l in pathlib.Path(sys.argv[1]).read_text().splitlines()]
fp, attempts = rows[0]["_fingerprint"], rows[1:]
passed = sum(a["passed"] for a in attempts)
tokens = sum((a["tokens_in"] or 0) + (a["tokens_out"] or 0) for a in attempts)
print(json.dumps({
"run_id": fp["run_id"],
"harness": fp["harness_sha256"],
"tasks": len(attempts),
"pass_rate": round(passed / len(attempts), 3),
"tokens_per_solved_task": round(tokens / passed, 1) if passed else None,
"p50_wall_clock_s": sorted(a["wall_clock_s"] for a in attempts)[len(attempts) // 2],
}, indent=2))
Now the part that keeps you honest. With 20 tasks, a pass rate near 50% has a standard error around 0.11, so a 95% interval spans roughly plus or minus 22 points. If your number differs from the published one by five points, you have measured nothing. You have re-measured noise with extra steps.
Step 5 — Diff the fingerprint, not the headline
When your result does diverge, compare fingerprints field by field before you compare conclusions.
| Your observation | Most likely cause | What to do next |
|---|---|---|
| Delta under ~10 points at n=20 | Sampling noise | Add tasks; do not publish a claim |
| Delta 25+ points | Different conditions | Diff fingerprint fields one at a time |
| Budget or retries differ | Not the same benchmark | Re-run both under one config |
| Task list unavailable | Unverifiable source | Keep it as directional context only |
| Verify command fails on the clean checkout | Flaky tests, not agent failure | Run verify 3x pre-patch, then exclude |
That last row matters more than it looks. Before blaming an agent, run each task's verify command on the unmodified checkout three times. A test suite that fails 1 in 3 times on clean code will manufacture agent failures out of nothing, and you will spend a weekend debugging a model that was fine.
Limitations, and who should not do this
- A free tier is a moving target. If the terms change mid-study, your run is invalid, not merely stale. Pin what you verified and re-run after a change.
- Wall-clock numbers from free, shared infrastructure are not latency benchmarks. Use them for timeout tuning only.
- 20 tasks cannot resolve a two-point difference. Do not pretend otherwise.
- Token accounting is only as good as the provider's usage reporting. If a field is
null, the summary showsnull.
Skip this approach if you need procurement-grade or regulatory-grade evidence, if you are comparing models that differ by under a few points, or if your organization cannot run untrusted agent-generated patches in an isolated worktree. That last constraint is real: an eval harness executes code a model wrote.
The short version
Fingerprint the claim. Pin the config. Let the verify command, not the harness, decide pass or fail. Record a ledger. Reduce it in code. Check the noise floor before you check your ego. Then ship the harness alongside the number, so the next person can re-derive it instead of trusting it.
If you want the shortest path to a first run, start from the MonkeyCode docs to confirm what free access currently includes, then paste the config above and swap in your own tasks — the docs are the only current source for the availability claims, and the config is the checklist to verify them against.
Top comments (0)