Your teammate drops a screenshot into the channel just after lunch. Twelve tasks. Twelve green checks. The caption asks you to paste the percentage into the weekly update.
You open the suite before you touch the slide. Three prompts name the function the hidden test calls. One fixture file still contains a comment that states the expected return value. Another task is a near-copy of a public tutorial the model has likely seen in training.
A perfect score, in that shape, is not a result. It is a harness smell. This walkthrough shows you how to add a negative-control slice so a headline number cannot hide leakage, a broken assertion, or a task the agent should have refused.
What a single percentage hides
A coding-agent score is a ratio. Numerator: trials you marked pass. Denominator: trials you chose to count. Change either side and the percentage moves, even if the model did not.
Public write-ups keep circling the same temptation. Someone reports zero mistakes on a small private set, and the number travels faster than the method. You do not need to litigate that post. You need a rule for your own suite: no solvable-slice pass rate gets quoted until the negative-control slice has been scored beside it.
That rule is methodological. It is not a product claim, and it is not a leaderboard.
The three slices you will actually run
Keep the suite small enough to rerun. Eight to sixteen tasks is enough to debug the harness. It is not enough to rank vendors. Label that limit in the report so nobody promotes the figure later.
Use three slices, stored as data, not as a vibe.
- Solvable. A real bug or a small feature. The prompt does not contain the patch, the test name, or the expected output.
- Negative control. A task the agent should not pass. Examples: the repo already matches the spec, the requested file does not exist, or the instruction contradicts the test. A pass here is a false pass.
- Canary. The same shape as a solvable task, plus a unique string that must never appear in a correct patch. If the model emits the canary, treat the trial as contaminated, not as a success.
You score the slices separately. You do not average them into one marketing percentage.
Step 1: Freeze the task card
Write one JSON object per task before you call any model. The card is the dataset. If a field is missing, the checker rejects the task instead of guessing.
{
"id": "neg-03",
"slice": "negative",
"prompt": "Add a discount function. Do not modify billing.py.",
"forbidden_paths": ["billing.py"],
"hidden_test": "tests/test_billing.py",
"canary": null,
"expected": "fail_or_abstain"
}
A solvable card sets "slice": "solvable" and "expected": "pass". A canary card sets "slice": "canary" and puts a random token in "canary", for example CANARY-7f3a9c. That token must not be a word the solution needs.
Do not generate these cards from the same chat that will later solve them. If you do, you have already leaked the rubric into the context you are grading.
Step 2: Pin the run controls
Controls are the boring part. They are also the part that stops a number from becoming an ad.
| Control | What you freeze | Why it matters |
|---|---|---|
| Attempt budget | Same max tries per task | Extra retries inflate pass rate |
| Tool list | Same editor, shell, test command | A richer tool schema is a different agent |
| Context | Prompt plus repo snapshot only | Hidden tests stay out of the prompt |
| Clock | One timeout value | A hung run is a fail, not a missing row |
| Scorer | This checker, versioned | Hand edits after the fact are a new study |
Write the freeze into run_envelope.json next to the tasks. If you change the timeout, you start a new run id. You do not edit yesterday's pass rate in place.
Step 3: Run the checker, not a vibe
The script below is a proposed checker. It has not been executed against a live model for this article. Treat it as a starting artifact. Run it yourself before you trust a row.
#!/usr/bin/env python3
"""Slice scorer for a tiny coding-agent suite. Unexecuted example."""
import argparse, json, sys
def load_jsonl(path):
with open(path, encoding="utf-8") as fh:
return [json.loads(line) for line in fh if line.strip()]
def false_pass(task, trial):
if task["slice"] != "negative":
return False
return trial.get("verdict") == "pass"
def canary_hit(task, trial):
token = task.get("canary")
if not token:
return False
blob = trial.get("patch", "") + trial.get("stdout", "")
return token in blob
def main():
parser = argparse.ArgumentParser()
parser.add_argument("--tasks", required=True)
parser.add_argument("--trials", required=True)
args = parser.parse_args()
tasks = {row["id"]: row for row in load_jsonl(args.tasks)}
counts = {}
for trial in load_jsonl(args.trials):
task = tasks[trial["task_id"]]
key = task["slice"]
bucket = counts.setdefault(
key,
{"n": 0, "pass": 0, "false_pass": 0, "canary": 0, "abstain": 0},
)
bucket["n"] += 1
if trial.get("verdict") == "pass":
bucket["pass"] += 1
if trial.get("verdict") == "abstain":
bucket["abstain"] += 1
if false_pass(task, trial):
bucket["false_pass"] += 1
if canary_hit(task, trial):
bucket["canary"] += 1
solvable = counts.get("solvable", {})
negative = counts.get("negative", {})
canary = counts.get("canary", {})
quote_ok = (
negative.get("false_pass", 0) == 0
and canary.get("canary", 0) == 0
and solvable.get("n", 0) > 0
)
report = {"slices": counts, "headline_allowed": quote_ok}
json.dump(report, sys.stdout, indent=2)
print()
if __name__ == "__main__":
main()
Run it like this after your batch finishes:
python3 slice_scorer.py --tasks tasks.jsonl --trials trials.jsonl > report.json
python3 -m json.tool report.json
headline_allowed stays false when any negative-control trial passes, or when any canary token shows up in a patch or log. That is the point. A green solvable slice cannot outvote a dirty control slice.
Step 4: Read the metrics in the right order
Look at four numbers, in this order, and stop if an earlier one is bad.
- Canary hit count. Any hit means the prompt, the log, or the scorer leaked a marker into the artifact. Fix the harness. Do not quote a pass rate.
- False-pass count on the negative slice. A pass means the tests are weak, the agent ignored a constraint, or you labeled a solvable task as negative. All three are measurement bugs.
- Abstain rate. An agent that never abstains on impossible tasks is not more helpful. It is willing to invent a diff. Record abstains as their own outcome. Do not fold them into failures without saying so.
-
Solvable pass rate. Only now. Report it as
passes / solvable_nfor this run id. Attach the envelope file. Do not call it accuracy, and do not compare it with a run that used a different retry budget.
If you want a sentence for the weekly update, use this shape: "On this run id, solvable pass rate was X/Y after zero false passes and zero canary hits. N is too small to rank models." Fill X and Y from your report. Leave them blank until the checker prints them.
Where a free run fits
You can execute this loop on a laptop. You can also park the batch on a free server so a dropped SSH session does not eat the run, and call a model through free access so you are not burning a paid quota while you debug the harness.
Disclosure: This article was prepared as part of MonkeyCode's product outreach.
MonkeyCode is the open-source project this draft was asked to cover. The operator supplied two availability claims only: free model access, and a free server option. This article does not name a model, a token allowance, a machine size, or a duration. Those details change. Read the current project docs before you plan a batch, and do not cite a quota from memory or from an older post.
Use the free path for harness debugging, not for a launch number. Generate candidate patches with whatever free model access you currently have. Store each trial as one JSONL row: task_id, verdict, patch, stdout, attempt. Run slice_scorer.py on the server, or copy the JSONL home and run it locally. Keep the raw file. A percentage without the file is a rumor.
If you do not have that access, nothing in the method depends on it. The checker is plain Python. The discipline is the product of the method, not of the host.
Why this number is not marketing
Marketing wants one figure, a rising arrow, and a model name in the title. This method refuses that shape on purpose.
A solvable pass rate with a dirty negative slice is compatible with a broken test, a leaked answer, or a lucky retry. Publishing it as "the agent scores 100%" trains your readers to ignore the denominator. Separating slices makes the failure visible: the agent looked strong because the suite could not say no.
You also refuse cross-run averages. Eight tasks on Tuesday and eight different tasks on Friday are not a sixteen-task study. They are two anecdotes. Say so in the report footer.
Limitations
This slice will not save a bad dataset. If every negative control is obviously impossible, a model can fail them and still be weak on real work. If your canary is a common word, you will flag innocent patches.
Eight tasks cannot support a ranking, a confidence interval you would show a customer, or a claim about coding ability. The checker does not compile code. You still need a real test command in the runner that produces verdict. No model was scored for this draft, so there is no measured pass rate to cite.
Timeouts, flaky tests, and network calls will still move verdict. Pin them in the envelope, or accept that the score jitters.
Who should not use this
Skip the method if you need a statistically powered comparison. You want a larger, pre-registered suite and a frozen hidden set, not a lunch-sized slice.
Skip it if the code is safety-critical. A JSONL checker is not a security review, a license review, or a substitute for reading the diff.
Skip it if you already plan to quote a single percentage in an ad. The checker is designed to withhold that sentence. Fighting it means you wanted a slogan, not a measurement.
What you do next
Pick four solvable tasks from a repo you maintain. Write four negative controls that should not pass. Add two canaries. Freeze the envelope. Run the checker. If headline_allowed is false, fix the suite before you mention a model.
When you want a place to park that first noisy batch, a free MonkeyCode server session is enough to keep the JSONL intact while you iterate. Confirm the current limits yourself, then archive the report next to the tasks. The archive is the result. The percentage is optional.
Top comments (0)