I will not publish best-of-five. A free code model is not usable because one sample passed. That sample might be the only one that would have passed. If you publish the highlight reel, you measured a slot machine, not a tool.
You have seen the flattering cut, haven't you? Five draws, one green check, a screenshot, a conclusion. What question did that answer? Not the question your teammate hits on a Tuesday, when nobody is sitting there rerolling until the tests blush.
The feed this week is busy with profile toys and portfolios people vibe-coded over a weekend. Fine toys. None of them tell you whether draw three still exits zero.
So the conclusion comes first, before any endpoint. Report the first draw, the stable pass, and the flip. Ignore any protocol that lets you delete the red rows. If those three numbers are not on disk, you do not have an evaluation. You have a mood.
First-draw pass is what an unattended user gets. Stable pass is the share of tasks that pass on every draw, with no runner error hiding in the set. Flip is the share that both pass and fail across draws.
Same prompt. Same check. Different verdict. That last one is the alibi killer.
A task that flips is not basically working. It is a coin, and coins do not get merge rights. Would you hand a coin your migration?
Why not a single accuracy number? Because a single number is how luck borrows a lab coat. Best-of-five can rise while stability falls. You would feel better and ship worse.
I would rather see an ugly stable-fail than a pretty flip. An ugly fail tells me to change the task class or the tool. A flip tells me I was about to quote a coincidence.
I do not have a leaderboard, and I am not going to invent one. This note is a protocol and a scorer. Any rows that look like results in the fixture are a self-check of the arithmetic. They are not a benchmark of a free model, a paid model, or a server.
If a pass rate shows up here without a CSV from the day the checks ran, treat it as fiction. I want you suspicious. Suspicion is the method.
Where do the free generator and the free runner enter? I want both boring and available, so the afternoon is about variance instead of about whether I could afford the draws. MonkeyCode is the option this harness is aimed at: free model access for the completions, and a free server option so every sample faces the same check command rather than whatever my laptop happened to have installed. Disclosure: This article was prepared as part of MonkeyCode's product outreach.
The outreach brief calls the project open source. Clone the repository you can verify. A paraphrased client is not a source.
I am also not quoting a token quota, a machine size, or a duration. Those are dashboard facts, and dashboard facts rot faster than blog posts. On the day you run, read the current allowance and the current server terms, then copy them into the run notes by hand.
If the allowance is smaller than K times your task count, shrink K before you start. Dropping tasks afterward, and still calling the file a replicate, is how samples vanish while the chart stays confident.
The harness is plain Python on purpose. No framework to debug before the model. It shells out to your check command, which is the whole point. Pretty code that never executes is a different hobby.
#!/usr/bin/env python3
"""split_score.py — first draw, stable pass, and flip.
Live mode is unexecuted until MODEL_ENDPOINT is set.
`--fixture` checks the arithmetic only. It is not a model score.
"""
from __future__ import annotations
import argparse, csv, json, os, subprocess, hashlib
from pathlib import Path
def load_tasks(root: Path) -> list[dict]:
tasks = []
for path in sorted(root.glob("*/task.json")):
task = json.loads(path.read_text())
task["dir"] = str(path.parent)
tasks.append(task)
if not tasks:
raise SystemExit(f"no tasks under {root}")
return tasks
def call_model(endpoint: str, prompt: str, draw: int) -> str:
# Proposal only. Match the real request shape before a live run.
# Read any key from the environment. Never hardcode one.
import urllib.request
body = json.dumps({
"prompt": prompt,
"draw": draw,
"temperature": 0.2,
}).encode()
req = urllib.request.Request(
endpoint,
data=body,
headers={"Content-Type": "application/json"},
)
with urllib.request.urlopen(req, timeout=60) as resp:
payload = json.loads(resp.read().decode())
if "text" not in payload:
raise SystemExit("response missing text; fix the client, do not score it")
return payload["text"]
def run_check(task_dir: str, code: str) -> str:
snippet = Path(task_dir) / "_candidate.py"
snippet.write_text(code)
try:
proc = subprocess.run(
["python3", "check.py"],
cwd=task_dir,
capture_output=True,
text=True,
timeout=30,
)
except subprocess.TimeoutExpired:
return "runner_error"
if proc.returncode == 0:
return "pass"
if proc.returncode < 0:
return "runner_error"
err = (proc.stderr or "") + (proc.stdout or "")
if "Traceback" in err and "_candidate" not in err:
return "runner_error"
return "fail"
def summarize(rows: list[dict]) -> dict:
by_task: dict[str, list[tuple[int, str]]] = {}
for row in rows:
by_task.setdefault(row["task"], []).append((int(row["draw"]), row["verdict"]))
n = len(by_task)
if n == 0:
raise SystemExit("no rows to score")
first = stable = flip = runner = 0
for pairs in by_task.values():
verdicts = [v for _, v in sorted(pairs)]
if any(v == "runner_error" for v in verdicts):
runner += 1
if verdicts[0] == "pass":
first += 1
if all(v == "pass" for v in verdicts):
stable += 1
if "pass" in verdicts and "fail" in verdicts:
flip += 1
return {
"tasks": n,
"first_pass": round(first / n, 3),
"stable_pass": round(stable / n, 3),
"flip_rate": round(flip / n, 3),
"tasks_with_runner_error": runner,
}
def main() -> None:
parser = argparse.ArgumentParser()
parser.add_argument("--tasks", default="tasks")
parser.add_argument("--draws", type=int, default=5)
parser.add_argument("--fixture", action="store_true")
parser.add_argument("--out", default="split_score.csv")
args = parser.parse_args()
if args.fixture:
# Synthetic rows. They test the scorer. They do not describe a product.
rows = [
{"task": "parse_int", "draw": 0, "verdict": "pass"},
{"task": "parse_int", "draw": 1, "verdict": "pass"},
{"task": "parse_int", "draw": 2, "verdict": "fail"},
{"task": "slug", "draw": 0, "verdict": "fail"},
{"task": "slug", "draw": 1, "verdict": "fail"},
{"task": "slug", "draw": 2, "verdict": "fail"},
{"task": "clamp", "draw": 0, "verdict": "pass"},
{"task": "clamp", "draw": 1, "verdict": "pass"},
{"task": "clamp", "draw": 2, "verdict": "pass"},
]
else:
endpoint = os.environ["MODEL_ENDPOINT"]
rows = []
for task in load_tasks(Path(args.tasks)):
for draw in range(args.draws):
code = call_model(endpoint, task["prompt"], draw)
verdict = run_check(task["dir"], code)
rows.append({"task": task["id"], "draw": draw, "verdict": verdict})
summary = summarize(rows)
with open(args.out, "w", newline="") as fh:
writer = csv.DictWriter(fh, fieldnames=["task", "draw", "verdict"])
writer.writeheader()
writer.writerows(rows)
digest = hashlib.sha256(Path(args.out).read_bytes()).hexdigest()[:12]
print(json.dumps({**summary, "csv_sha256_12": digest}, indent=2))
if __name__ == "__main__":
main()
The fixture is supposed to be mean. parse_int flips. slug fails every draw. clamp stays green.
If you cannot say those three sentences without opening a pricing page, you are not ready for a live file. What should the scorer print? First-draw 0.667, stable pass 0.333, flip rate 0.333, runner errors 0. I did that division on paper before I trusted a script. You should do the same.
A scorer you have not caught is just another lucky sample, wearing a .py hat. Run the self-check before you touch a network. It should not need a key, a server, or a hope.
python3 split_score.py --fixture --out /tmp/split_fixture.csv
python3 - << 'PY'
import json, subprocess
raw = subprocess.check_output(
["python3", "split_score.py", "--fixture", "--out", "/tmp/split_fixture.csv"]
)
got = json.loads(raw)
expect = {
"first_pass": 0.667,
"stable_pass": 0.333,
"flip_rate": 0.333,
"tasks_with_runner_error": 0,
}
for key, val in expect.items():
assert got[key] == val, (key, got[key], val)
print("scorer ok", got["csv_sha256_12"])
PY
A task is a directory, not a vibe. task.json holds an id and a prompt. check.py imports the candidate and exits non-zero when the behavior is wrong. Keep the check dumber than the task.
If the check is the clever part, you are grading your harness and calling it a model. That is an easy way to feel rigorous while learning nothing.
{
"id": "clamp",
"prompt": "Write clamp(x, lo, hi) in _candidate.py. Return x bounded to [lo, hi]. No imports. No prints."
}
# tasks/clamp/check.py — example check, not a live result
import _candidate
def expect(got, want):
if got != want:
raise SystemExit(f"{got!r} != {want!r}")
expect(_candidate.clamp(3, 0, 2), 2)
expect(_candidate.clamp(-1, 0, 2), 0)
expect(_candidate.clamp(1, 0, 2), 1)
The generator and the runner should not share a mood. Draws come from the free model endpoint. The live scorer runs on the free server, so check.py sees one image, one Python, one clock. Laptop runs are for the fixture. Mix them and you are grading your own machine again.
When the fixture matches, and only then, point the script at a live endpoint. The request body inside call_model is a proposal. I have not executed it against a current API, and I will not guess a path, a model id, or a header name in public.
Match the client to the docs you verified today. Keep the key in the environment. Log a hash of the prompt next to the CSV.
If the prompt changes, start a new file. Editing the old file is not replication. It is a quiet rewrite.
# Proposal, not a verified hostname or client path.
# Confirm today's free-server login from the dashboard, then run the scorer there.
export MODEL_ENDPOINT="https://example.invalid/draw"
python3 split_score.py --tasks ./tasks --draws 5 --out runs/today.csv
sha256sum runs/today.csv
Read the printout like someone who has been fooled by a green cell before. High first-draw plus high flip means the generator can look competent and still be a coin. Would you let that coin edit a migration? I would not.
Low first-draw and a stable fail is a kindness. It tells you this task class is a break, not a teaser. High stable pass and a flip near zero is the only pattern I will call performs well, and only for drafts a human still reads.
Runner errors are not model fails. They mean the check, the timeout, or the server session broke. Fix that, rerun, and do not average across the mess.
Who should walk away? If you need a number for a launch note tonight, walk away. Five draws times a dozen tasks will not fit in a coffee break, and a partial grid is not a result you get to round up.
If the code can touch credentials, payments, or auth, walk away from the free path. A green clamp says nothing about a webhook. If the check needs a private network, a production token, or customer data, do not place it on a free server either. Shared runners are for toy tasks with toy fixtures.
If your question is taste in prose, this harness will lie politely, because it only speaks exit codes. Exit codes do not have taste. They have zero and not-zero.
There is a quieter limit, and it matters. Temperature 0.2 in the proposal is a frozen confound, not an optimum I measured. Change it and you have a different study. K=5 is a budget, not a proof about sample size.
A free server can restart, throttle, or run slower than a laptop. Those events belong in the notes, not in the model column. I will not use this protocol to rank products against each other.
It can rank task classes against one generator, on one day, under one check suite. That is enough. It is also all the certainty the design earns.
If you already have a MonkeyCode login, point the draws at the free model, run the checks on the free server, and keep the CSV beside the prompt hash. If you do not, run the fixture anyway. The arithmetic does not need a brand. The brand does not get a score until the arithmetic has a file.
Top comments (0)