DEV Community

Casey Zhang
Casey Zhang

Posted on

Name the Failure Class Before You Rank a Coding Agent

You are halfway through a review when a teammate drops a screenshot in the channel. The caption is a coding-agent pass rate. The log is a scroll away, and almost nobody opens it.

You do. Compile failures sit beside timeouts. A green row has edited a file the task marked read-only. The caption is still up in the channel. The evidence is not in the caption.

This piece is a method for that moment. You will keep a failure-class ledger beside every coding-agent score you intend to cite. The ledger names how each task ended, and it includes a negative-control slice that a correct run must not pass. If you cannot publish the ledger, you do not publish the percentage.

What a blended pass rate conceals

A coding-agent run is not a coin flip. The agent can fail to compile, fail a test, time out, return no patch, or wander into files it was told to leave alone. Those endings are different events.

A single green rate treats a timeout like a near miss and a forbidden edit like a success. You already know a score needs a frozen run envelope before the percentage is worth quoting. This workflow adds the next artifact: a class for every row, counted before anyone ranks a model.

The classes belong in the dataset contract. They are not a story you attach after the chart looks good. Public write-ups still open on one figure. You can notice that habit without borrowing their numbers, their tasks, or their captions.

The ledger you are going to build

You need three files, and you need them before the call.

  1. A task file that marks each item positive or negative.
  2. A run log in JSONL, one object per finished task, appended by the harness.
  3. A classifier that refuses a citeable rate when the controls fail.

No model name belongs on the cite line until those three exist. The script below is a proposed harness. It is not a measured leaderboard. The sample rows are illustrative, so you can see the rule fire. They are not results from a named model, and they are not a claim about any release this week.

Dataset, metrics, and controls

Dataset. Split the set before you call any model. Positive tasks have a known acceptable patch and a test command that can fail. Negative-control tasks have no acceptable patch. That can be a planted failure in a file the agent must not touch, a test that must stay red, or a prompt that asks for a change outside the allowed paths.

Keep a real negative slice, not a single toy row. A set of only solvable tasks cannot detect a harness that marks everything pass. Date the task file. If you later suspect a row has leaked into public training text, retire it in the notes and say so. Do not swap in an easier task and keep the old chart.

Metrics. Report counts, then rates. Use the positive slice only for the pass rate. Use the negative slice only for the false-pass rate. Across the whole run, also report compile-fail share, test-fail share, timeout share, over-edit share, and unresolved share. Unresolved means a missing log row, a crash, or an outcome the schema does not name.

Controls. Freeze these before the first call, and write them into a header the classifier can see.

  • Harness version and prompt-template hash.
  • Timeout, retry cap, and token budget, identical for every row you will compare.
  • Allowed path list per task.
  • Seed list. One seed is a demo. Three seeds is the minimum if you intend to cite a rate.
  • A cite rule, written in advance, for when the rate may be quoted at all.

Do not drop a class because it makes the chart uglier. That drop is how a method turns into a slogan. A number you can only publish after deleting rows is marketing, even if the arithmetic is correct.

How to build a negative control that is not a trick

A negative control is a task the agent must not solve. You are not trying to embarrass a model. You are checking whether the harness awards a pass for the wrong reason.

Three shapes work on ordinary application code you own.

  1. A read-only sentinel. The prompt asks for a feature under src/. A known bad line lives in a file the task marks forbidden, and the grader fails if that file's hash changes. If the agent edits it and tests go green, the outcome is over_edit, not pass.
  2. An unsatisfiable local test. The task asks for a function to return two incompatible results for the same input, and you actually run the checker. A pass means the grader did not run, or someone weakened the test.
  3. A planted red test outside the allowed diff. The agent may edit billing.py only. A second test is already failing and must stay failing. If the run deletes that test, log over_edit.

Write the expected class in the task file before the call. If you invent the expected class after you see the patch, you are narrating. You are not measuring. Keep the control inside a repository you can explain to a reviewer. This is a grader check, not a security exercise.

Step by step: build the classifier

Work on a small box you can wipe. The ledger itself is a file and a short script. A free server can hold both. Model calls are a separate budget. Do not assume a tier, a quota, or a machine shape you have not read on the day you run.

1. Fix the outcome schema

Create a short schema note in the run directory and treat it as code. Allowed outcome values are pass, compile_fail, test_fail, timeout, no_patch, over_edit, and unresolved.

If the agent edits a forbidden path, the outcome is over_edit even when tests are green. A green test on a negative-control task stays pass in the raw log. The classifier, not the agent, turns that row into a false pass. You want the raw outcome honest. You want the judgment in the script, where a reviewer can see it.

2. Append one JSON object per task

Have the harness append, not overwrite. A crash should leave a partial file you can still score. Do not backfill missing rows from memory.

mkdir -p runs/ledger-demo
cat > runs/ledger-demo/header.json << 'EOF'
{
  "harness_version": "0.3.1",
  "prompt_hash": "sha256:replace-me",
  "timeout_s": 120,
  "max_retries": 0,
  "token_budget": 8000,
  "seeds": [7, 19, 23]
}
EOF
Enter fullscreen mode Exit fullscreen mode

The hash and the budget in that header are placeholders for the envelope you actually freeze. Replace them. Do not cite a run whose header is still the example. 8000 is a stand-in so the field exists. It is not a recommended product limit.

Illustrative rows, not an evaluation:

{"task_id":"pos-01","slice":"positive","outcome":"pass","seed":7}
{"task_id":"pos-02","slice":"positive","outcome":"compile_fail","seed":7}
{"task_id":"pos-03","slice":"positive","outcome":"timeout","seed":7}
{"task_id":"neg-01","slice":"negative","outcome":"test_fail","seed":7}
{"task_id":"neg-02","slice":"negative","outcome":"pass","seed":7}
Enter fullscreen mode Exit fullscreen mode

neg-02 is the row that should stop you. The log recorded a pass on a task with no acceptable fix. A caption would have buried that row inside a friendly rate.

3. Classify, then apply the cite rule

Save the following as failure_ledger.py. Run it on your own logs. The demo file exists so you can watch the rule reject a flattering mix. It is not evidence about a product.

#!/usr/bin/env python3
"""Failure-class ledger. Proposed harness, not a published score."""
import json
import sys
from collections import Counter

CLASSES = (
    "pass",
    "compile_fail",
    "test_fail",
    "timeout",
    "no_patch",
    "over_edit",
    "unresolved",
)

def load_jsonl(path):
    rows = []
    with open(path, encoding="utf-8") as handle:
        for line in handle:
            line = line.strip()
            if line:
                rows.append(json.loads(line))
    return rows

def citeable(n_pos, n_neg, unresolved, false_passes):
    # Pre-register this rule before you look at model identity.
    if n_pos == 0 or n_neg == 0:
        return False, "missing positive or negative slice"
    if unresolved:
        return False, "unresolved rows present"
    if false_passes:
        return False, "negative control recorded a pass"
    return True, "class ledger complete; quote counts beside the rate"

def main(path):
    rows = load_jsonl(path)
    by_class = Counter(r.get("outcome", "unresolved") for r in rows)
    n_pos = sum(1 for r in rows if r.get("slice") == "positive")
    n_neg = sum(1 for r in rows if r.get("slice") == "negative")
    pos_pass = sum(
        1
        for r in rows
        if r.get("slice") == "positive" and r.get("outcome") == "pass"
    )
    false_pass = sum(
        1
        for r in rows
        if r.get("slice") == "negative" and r.get("outcome") == "pass"
    )
    unresolved = sum(1 for r in rows if r.get("outcome") not in CLASSES)
    ok, reason = citeable(n_pos, n_neg, unresolved, false_pass)
    pos_rate = (pos_pass / n_pos) if n_pos else None
    print("rows", len(rows))
    print("positive_pass", pos_pass, "/", n_pos, "rate", pos_rate)
    print("false_pass", false_pass, "/", n_neg)
    for name in CLASSES:
        print("class." + name, by_class.get(name, 0))
    print("citeable", ok)
    print("reason", reason)

if __name__ == "__main__":
    if len(sys.argv) != 2:
        sys.stderr.write("usage: python3 failure_ledger.py runs/seed.jsonl\n")
        sys.exit(2)
    main(sys.argv[1])
Enter fullscreen mode Exit fullscreen mode

Run the illustrative file first:

python3 failure_ledger.py runs/ledger-demo/seed-7.jsonl
Enter fullscreen mode Exit fullscreen mode

On that file, citeable is false because a negative control passed. That is the method working. A caption-style summary would have counted green rows and moved on.

After you have real logs, score every seed the same way:

for seed in 7 19 23; do
  python3 failure_ledger.py "runs/ledger-demo/seed-${seed}.jsonl"
done
Enter fullscreen mode Exit fullscreen mode

Compare class counts across seeds before you compare models. If one seed is mostly timeouts and the next is mostly test failures, you do not have a stable rate yet. You have noise. Wait.

4. Read the decision table before you write the post

Condition What you may say What you must not say
Missing negative slice Harness smoke test only Any pass rate as a quality claim
Any false pass Negative control failed; do not rank Solved the set
Unresolved rows dropped from the denominator Nothing comparative A higher rate than a run that kept those rows
Classes published, false passes at zero, envelope frozen Positive pass rate with class counts and seeds A bare percentage
Budgets or timeouts differ across endpoints Separate logs, not a ranking One endpoint beat the other

The right column is the marketing failure mode. If your draft has the number and none of the conditions, delete the number. Readers can trust a smaller, qualified rate more than a clean one they cannot recompute.

5. Point the driver at the endpoint you actually have

The ledger does not care which vendor answered. It cares that every row shares the header you froze.

Disclosure: This article was prepared as part of MonkeyCode's product outreach. Free model access and a free server option can carry this loop when you do not already have a box and an endpoint: the server stores the JSONL and runs the classifier, and free model access is only the call path. This article does not state a quota, a hardware shape, a model list, a duration, or a promise that the free tier will stay as it is. Read the current plan before you depend on it. If that plan changes, keep the ledger and move the call path.

One check after the first seed: open the header and confirm the prompt hash, timeout, and token budget match the run you think you launched. A drifted template is a different experiment. Score it separately, or do not score it.

What to do when the first ledger is red

Your first real run will probably print citeable false. Leave that line in the notes. It means the rule survived contact with a log.

Change one control, not three.

  • If the reason is a false pass, inspect the grader before you blame the agent. A test that always exits zero is a harness bug.
  • If the reason is unresolved rows, fix the append path. A lost row is not a timeout. Do not reconstruct it.
  • If timeouts dominate, you may raise the timeout. That is a new envelope. Keep the old log. Do not overwrite it and pretend both runs shared a clock.

Rerun the same seed. Compare class counts, not impressions. Spend a second seed only after you can name the largest class without looking at the caption. A single flattering seed is how the screenshot in the opening gets made.

Comparing two endpoints without inventing a winner

You can point the same task file at two endpoints. Freeze the header, swap only the base URL, and write separate JSONL files. Rank them only when both ledgers are citeable and both headers match on timeout, retry cap, and token budget.

If one call path stops early because a limit cut the run, mark the missing rows unresolved. Do not impute passes. A shorter bill is not a better agent. Write the stop reason into the header so a reader can see the run ended for a budget reason, not a quality reason.

Keep the two logs in separate directories. Mixing them into one average is how unequal budgets sneak back into the chart. If you cannot show both headers, you cannot show a winner.

Why the ledger is not a slogan

A class ledger is a poor advertisement. It shows timeouts, compile failures, and false passes that one rate can hide. That is why you keep it.

You pre-register the cite rule in code, then you run. You do not edit the rule after you see which endpoint looks kinder. Publish the counts even when the positive pass rate is flattering. Readers can recompute.

A figure they cannot recompute is a claim about your editing, not about the agent. Link the harness commit and the task-file commit next to any rate. Without those two hashes, the number cannot be reproduced, so it should not be cited.

Limits, and who should skip this

This method does not measure whether a patch is maintainable, secure, or what a user wanted. It measures whether your harness can tell a real pass from a miss, and whether you will show the miss.

Skip it when any of these are true.

  • You cannot inspect the tasks, so you cannot build a negative slice you trust.
  • You need a ranking today and will quote the pass rate even if the script prints citeable false.
  • Your real question is human preference or production incident rate. Use a different instrument.

A tiny or obvious negative slice can lie in the other direction. A false-pass rate of zero means little if every control is trivial. Rotate controls when they become familiar in your prompt archive. This is also the wrong tool if you wanted a security study. Do not turn the negative slice into an attack exercise. Keep it inside code you own, with tests you can explain to someone who was not in the room.

Hand-built sets overfit the failures you remembered to name. Twenty tasks can teach you whether the harness is honest. They cannot crown a general coding agent. Say the sample size in the same sentence as the rate, or do not say the rate.

What you should ship with the score

Ship four files, not a screenshot.

  1. header.json with the frozen envelope.
  2. The JSONL log for every seed.
  3. The classifier output, including citeable and the reason.
  4. A short note naming who wrote the tasks and whether any row was removed after the run.

If those files are missing, keep the percentage in the draft and out of the post. The ledger is the result. The rate is a summary of one column, and only after the other columns are clean.

If you already have a free model endpoint and a small free server, run one seed of your own tasks through this script before you paste a percentage into a review thread. Keep the ledger next to the figure. That pairing is the part worth sharing.

Top comments (0)