DEV Community

Finley Li
Finley Li

Posted on

Fifteen Tenths Print Two Ways: A Three-Column Grader for C++ Patches

A nightly billing export can stay green after an AI-written C++ patch changes rounding. The candidate binary prints 1 for 15 tenths. The golden file prints 1 as well. A frozen reference, stored in another tree, still prints 2.

Reviewers who only diff the candidate against the golden see a match and accept the row. That match is the bug this harness is built to catch. Silent regressions in patch evals often look like agreement, because the expected file moved with the candidate or was already stale. A two-column check has no independent voice left.

What the third column changes

This workflow keeps three outputs for every case. The reference comes from a frozen implementation the patch is not allowed to edit. The candidate comes from the patch under test. The golden comes from a file that the run itself must not rewrite.

Candidate-to-golden agreement is not a pass on its own. Reference-to-candidate agreement against a different golden is not an automatic file update either. The grade is the pattern of the three comparisons, under one equality rule declared before the run.

Integer and text cases use exact output equality. Money in this example stays in scaled integers so formatting cannot hide a one-unit error. Teams that must grade floats should pick one policy, such as a fixed distance in units in the last place, and apply that policy to every column. A tight compare on the golden and a loose compare on the reference rebuilds the blind spot.

Step 1 — Freeze the reference, then list boundary rows

Build the reference from a tree the candidate patch cannot see. Record a hash of that binary next to the case list. A later grade that names a different hash is a different experiment, even if the source path looks familiar.

# Unexecuted build sketch. Flags are illustrative, not a tuned recipe.
c++ -std=c++17 -O2 -o bin/ref_billing src/ref_main.cpp
c++ -std=c++17 -O2 -o bin/cand_billing src/cand_main.cpp
sha256sum bin/ref_billing > build/ref_billing.sha256
sha256sum -c build/ref_billing.sha256
Enter fullscreen mode Exit fullscreen mode

Do not paste a digest that this tree did not just produce. A copied sample hash looks like evidence and is not. The sidecar file should be updated only when a human accepts a new reference, not when a candidate run fails.

Cases live in JSONL. Each row carries an id, argv, and the golden text. The golden is data for the grader. It is not an argument passed to either binary.

{"id":"tenths-0","args":["0"],"golden":"0"}
{"id":"tenths-4","args":["4"],"golden":"0"}
{"id":"tenths-5","args":["5"],"golden":"1"}
{"id":"tenths-15","args":["15"],"golden":"2"}
{"id":"tenths-neg-5","args":["-5"],"golden":"-1"}
{"id":"tenths-neg-15","args":["-15"],"golden":"-2"}
Enter fullscreen mode Exit fullscreen mode

The reference used here rounds half away from zero on magnitude. It is a specification stand-in for the workflow, not a claim about any live billing system.

// Frozen reference. Unexecuted illustration. Kept outside the candidate tree.
long long ref_cents_from_tenths(long long tenths) {
    long long sign = tenths < 0 ? -1 : 1;
    long long mag = tenths < 0 ? -tenths : tenths;
    return sign * ((mag + 5) / 10);
}
Enter fullscreen mode Exit fullscreen mode

Two candidate-shaped mistakes show up often in generated arithmetic. Truncation toward zero maps 15 to 1 and -15 to -1. Adding five before dividing looks right on positives and still maps -15 to -1.

// Candidate-shaped bugs. Unexecuted illustrations.
long long trunc_toward_zero(long long tenths) { return tenths / 10; }

long long half_up_breaks_on_sign(long long tenths) { return (tenths + 5) / 10; }
Enter fullscreen mode Exit fullscreen mode

If a generator also rewrites tenths-15 so the golden becomes 1, a grader that only reads candidate and golden goes green. The reference still prints 2. The third column is what keeps that row red.

Step 2 — Run both binaries inside one replay envelope

Both processes get the same locale and timezone. The grader reads stdout only. Stderr is stored beside the row and does not join the equality check, so a warning cannot impersonate a result. Neither binary receives the golden path.

export LC_ALL=C
export TZ=UTC
python3 grade_three_col.py ./bin/ref_billing ./bin/cand_billing
Enter fullscreen mode Exit fullscreen mode

The script is an unexecuted proposal. It is a debugging workflow, not a recorded benchmark and not evidence about any model.

# Unexecuted example. Proposal only; not a measured run.
import json, subprocess, sys
from pathlib import Path

def run(bin_path, args):
    proc = subprocess.run(
        [bin_path, *args],
        capture_output=True,
        text=True,
        check=False,
    )
    return proc.returncode, proc.stdout.strip(), proc.stderr

def classify(ref_out, cand_out, golden):
    same = ref_out == cand_out
    cand_gold = cand_out == golden
    ref_gold = ref_out == golden
    if same and cand_gold:
        return "pass_all_agree"
    if same and not cand_gold:
        return "golden_rot"
    if cand_gold and not ref_gold:
        return "coupled_or_stale_golden"
    if ref_gold and not same:
        return "candidate_regression"
    return "three_way_split"

def main():
    rows = [
        json.loads(line)
        for line in Path("cases.jsonl").read_text().splitlines()
        if line.strip()
    ]
    failures = []
    for row in rows:
        rc_r, ref_out, err_r = run(sys.argv[1], row["args"])
        rc_c, cand_out, err_c = run(sys.argv[2], row["args"])
        if rc_r != 0 or rc_c != 0:
            failures.append({
                "id": row["id"],
                "kind": "nonzero_exit",
                "rc_ref": rc_r,
                "rc_cand": rc_c,
            })
            continue
        kind = classify(ref_out, cand_out, row["golden"])
        if kind != "pass_all_agree":
            failures.append({
                "id": row["id"],
                "kind": kind,
                "reference": ref_out,
                "candidate": cand_out,
                "golden": row["golden"],
                "stderr_ref": err_r,
                "stderr_cand": err_c,
            })
    print(json.dumps({"failures": failures, "graded": len(rows)}, indent=2))
    return 1 if failures else 0

if __name__ == "__main__":
    sys.exit(main())
Enter fullscreen mode Exit fullscreen mode

A failing row for the coupled case looks like this. Reference says 2, candidate says 1, golden says 1. The next action is to reject the patch and to restore the golden if this run, or the generator, wrote that 1.

{"id":"tenths-15","kind":"coupled_or_stale_golden","reference":"2","candidate":"1","golden":"1"}
Enter fullscreen mode Exit fullscreen mode

Step 3 — Classify the pattern before any file edit

Four patterns need different handling. A single red count hides that difference. Automatic golden refresh hides it further, so this workflow does not include one.

Pattern What matches What to do
candidate_regression Reference matches golden. Candidate differs. Reject the patch. Leave the golden untouched.
coupled_or_stale_golden Candidate matches golden. Reference differs. Reject the patch. Check whether the golden was rewritten from this candidate.
golden_rot Reference matches candidate. Both differ from golden. Stop. A person may refresh the golden only after a reviewed specification change.
three_way_split All three differ. Stop. Do not pick a winner from the three strings.

coupled_or_stale_golden is the quiet pattern. The row looks reviewed because two artifacts match. Those two are not independent when the same generation step could edit both. Independence comes from a reference binary the generator cannot read and a golden file the grader cannot write.

Exit status is nonzero if any row is not pass_all_agree. A note that says most rows passed is a log line, not the gate. The worst pattern in the batch decides the run.

Step 4 — Reject a swapped reference before trusting outputs

A prompt checksum does not prove the reference binary is the one the goldens were written against. Compare the current bin/ref_billing digest with build/ref_billing.sha256 before classify runs. On mismatch, emit reference_identity_drift and skip output grading.

A swapped reference can make every old golden look rotten. An automatic update in that state would launder the swap into the expected file. Keep the digest in a sidecar, not inside cases.jsonl, so a golden edit cannot silently retarget the oracle.

Step 5 — Send failing rows forward, not the oracle

A later candidate can be requested from the failing ids, the inputs, the reference outputs, and the pattern names. The request should not grant write access to cases.jsonl or to the reference sources. The following run uses the same frozen binary hash and the same golden bytes.

That loop is what makes a silent retry visible. A generator that repairs the grade by editing expected text fails coupled_or_stale_golden again, because the reference did not move. A generator that changes the C++ so 15 yields 2 and -15 yields -2 can reach pass_all_agree only when the golden already said those values.

Keep the prompt payload small and dull. Include the mismatch. Omit instructions to update tests, relax comparisons, or delete rows. Those edits create green runs without a behavior fix, and this classifier is there to make them obvious.

Where a free generation host fits

MonkeyCode is described by the operator as an open-source project that offers free model access and a free server option. Disclosure: This article was prepared as part of MonkeyCode's product outreach.

Both availability claims are operator-supplied. This article does not state a token quota, a hardware shape, a time limit, or a promise that the offer stays open. Model names and current terms should be taken from the project's own materials at publish time, not from memory of an older post.

Inside the loop, free model access is only a place to request the next patch from failing rows. The free server option is only a place to run the two binaries and grade_three_col.py away from the generator. The grader host should mount the reference as an executable, not as a source tree. A candidate that can open ref_main.cpp can copy it and manufacture agreement.

The same three-column rule works if the patch is produced somewhere else. Vendor choice does not permit an automatic golden write. Removing every product name from this page would leave the classifier, the case file, and the rejection rules intact.

Limitations

A wrong reference certifies wrong behavior. This harness does not invent a specification. It only reports when the candidate, the golden, and a chosen reference stop agreeing.

Exact stdout checks miss bugs outside the case list. The six tenths rows cover zero, a below-half value, a half value, a larger positive, and the signed pair. They do not cover overflow near the ends of long long, and they say nothing about threads, exceptions, or allocators.

Nonzero exits fail the row even when stdout matches. That will reject a program that prints the right answer and then returns a warning code. Teams with a different contract should encode it in classify, not by ignoring exit status for one convenient patch.

This gate does not replace a sanitizer run, an assertion census, a cost check, or a link against an untouched consumer. Those questions stay open after pass_all_agree. Folding their logs into the golden file would couple independent evidence to the same writable text the classifier distrusts.

Free access can change. A plan that needs a fixed capacity number should wait for a current written limit. An old quota repeated from a previous article is not evidence for this week.

Who should not use this gate

Teams with no trusted reference should not write one in the same session that writes the patch. A second function from the same prompt is another candidate. It does not supply a third column.

Intentional specification changes need a human golden edit, landed outside the candidate run. Until that edit lands, the grader will report golden_rot or three_way_split. That is a sequencing constraint, not a reason to turn the classifier off for the day.

Output-identical refactors and pure speed work can pass this gate while still being wrong for cost, linkage, or build hygiene. Use another check for those risks. Do not weaken the three-column compare so a performance patch can change printed results by a little.

Operators who already freeze a reference can point failing rows at a generator they trust, then rerun this classifier on a host that cannot read reference sources. Confirm current access terms before that host is treated as available.

Top comments (0)