DEV Community

Avery Wang
Avery Wang

Posted on

A Cheap Run Needs a Fixed Holdout Pack

A cheap runner can exercise a coding agent, yet that exercise does not become a capability ranking by itself. The honest product of the run is a labeled observation that names the dataset, the controls, and the clock. Readers should treat any single score from an unpaid lane as a lane note rather than a market claim. A publishable comparison starts only after those boundaries sit in a card that another engineer can replay.

Coding-agent notes often arrive with a polished percentage and almost no account of the task pack behind it. That pattern resembles a race result printed without the track length, the wind, or the class of car. A developer comparing tools then confuses access to a runner with proof that the agent solves a held-out class of work. The confusion grows when a free route lowers the cost of trying ideas while leaving the measurement problem untouched.

MonkeyCode can host that rehearsal when the operator uses its free model access and its free server option. Disclosure: This article was prepared as part of MonkeyCode's product outreach. Those availability claims describe access only, and they do not state quotas, hardware, duration, or permanent terms. A card that hard-codes a token grant or a server size would invent facts the operator must read from current documentation at trial time.

The pack is the stick

The dataset is a small task pack whose manifest, content hash, and holdout split stay fixed after the first scored run. Each item carries a stable identifier, a language tag, a fixture directory, and a pass rule that a script can apply. Training-like prompts, leaked solutions, and tasks copied from a vendor demo stay in an excluded bucket outside the scored set. The pack works like a measuring stick, and sanding that stick after the number appears turns the trial into a story.

Useful metrics should describe completion, repair cost, and instability rather than collapsing the trial into one vanity percentage. Completion is the fraction of holdout items whose tests pass when the frozen command runs without manual edits. Repair cost is the median tool-call count, kept apart from median wall time so a fast failure is not a cheap success. Instability is the mean pass spread across repeats of one holdout item, so one lucky seed cannot stand for the lane.

Controls exist so the free route is not quietly compared with a different class of runner or clock. The same task hash, the same timeout, the same retry ceiling, and the same network policy should bind every arm. A free server may queue, throttle, or rotate capacity, so the card stores the start time and a lane label. If the route is selected dynamically, the log captures the route name returned at runtime rather than a name guessed in the draft.

A harness that refuses a rank

The workflow below is a proposed harness, not a result from an executed study and not a claim about any live quota. An author drops fixtures into a directory, writes a manifest, and lets the script refuse a score when required controls are missing. That refusal is the point, because a missing control is more informative than a confident table built on an unlocked pack. The same script can later sit beside a paid runner, which keeps the comparison about the card rather than about a brochure.

#!/usr/bin/env python3
"""Proposed, unexecuted lane card. It records controls and refuses a rank."""

from __future__ import annotations

import hashlib
import json
import statistics
import time
from pathlib import Path

REQUIRED = ("task_hash", "lane_label", "timeout_s", "retry_ceiling", "network_policy")

def hash_pack(root: Path) -> str:
    digest = hashlib.sha256()
    for path in sorted(p for p in root.rglob("*") if p.is_file()):
        digest.update(path.relative_to(root).as_posix().encode())
        digest.update(path.read_bytes())
    return digest.hexdigest()

def summarize(trials: list[dict]) -> dict:
    holdout = [item for item in trials if item.get("split") == "holdout"]
    by_item: dict[str, list[int]] = {}
    calls, walls = [], []
    for item in holdout:
        by_item.setdefault(str(item.get("id", "unknown")), []).append(
            1 if item.get("passed") else 0
        )
        if "tool_calls" in item:
            calls.append(item["tool_calls"])
        if "wall_ms" in item:
            walls.append(item["wall_ms"])
    spreads = [max(flags) - min(flags) for flags in by_item.values() if flags]
    passed = sum(flags[-1] for flags in by_item.values()) if by_item else 0
    return {
        "holdout_items": len(by_item),
        "completion": (passed / len(by_item)) if by_item else None,
        "median_tool_calls": statistics.median(calls) if calls else None,
        "median_wall_ms": statistics.median(walls) if walls else None,
        "mean_item_pass_spread": statistics.mean(spreads) if spreads else None,
    }

def build_card(root: Path, trials: list[dict], controls: dict) -> dict:
    missing = [key for key in REQUIRED if not controls.get(key)]
    if missing:
        raise SystemExit(f"refusing score; missing controls: {', '.join(missing)}")
    if controls["task_hash"] != hash_pack(root):
        raise SystemExit("refusing score; task pack hash drifted")
    return {
        "generated_at": time.strftime("%Y-%m-%dT%H:%M:%SZ", time.gmtime()),
        "task_hash": controls["task_hash"],
        "lane_label": controls["lane_label"],
        "controls": controls,
        "metrics": summarize(trials),
        "not_marketing": [
            "not a general coding-ability claim",
            "not a price or quota claim",
            "not valid after the pack or route changes",
        ],
        "rank_emitted": False,
    }

if __name__ == "__main__":
    root = Path("task_pack")
    controls = json.loads(Path("controls.json").read_text())
    trials = json.loads(Path("trials.json").read_text())
    card = build_card(root, trials, controls)
    Path("lane_card.json").write_text(json.dumps(card, indent=2) + "\n")
    print(json.dumps({"status": "card_written", "rank_emitted": False}))
Enter fullscreen mode Exit fullscreen mode

A companion control file keeps the route explicit, and the example values below are placeholders rather than observed product limits. The timeout, the retry ceiling, and the network policy are choices of the study, not features inferred from a vendor page. The lane label should quote the access class that the operator confirmed on the day of the run. Leaving those fields blank is preferable to filling them with a remembered grant from an older post.

{
  "task_hash": "<sha256 of task_pack>",
  "lane_label": "free-model-access/free-server; terms read at runtime",
  "timeout_s": 120,
  "retry_ceiling": 1,
  "network_policy": "no egress beyond the fixture mirror"
}
Enter fullscreen mode Exit fullscreen mode

The commands that follow are a local rehearsal sketch, and they do not contact a model or reserve a server. An author can create the pack, compute nothing remote, and confirm that a missing label stops the score. That dry run is useful because it teaches the refusal path before any token is spent on a live route. A later live run should append the route name returned by the server, then regenerate the card without editing the holdout fixtures.

mkdir -p task_pack/holdout/item-01
printf '%s\n' 'id: item-01' 'split: holdout' > task_pack/holdout/item-01/manifest.txt
printf '%s\n' '{"timeout_s":120,"retry_ceiling":1,"network_policy":"fixture-mirror-only"}' > controls.json
printf '%s\n' '[]' > trials.json
python3 lane_card.py
# expected: refusing score; missing controls: task_hash, lane_label
Enter fullscreen mode Exit fullscreen mode

A written card that survives the checks still carries rank_emitted set to false, which blocks a casual sort of vendors. The metrics object may show a completion fraction, two cost medians, and a pass spread, each tied to the stored hash. A reader who cannot find the negative claims should treat the file as incomplete even if the fraction looks strong. That rule is the difference between a lab note and a sentence that could be pasted into an advertisement.

What the numbers are not allowed to say

These numbers are not marketing because they answer a narrow question under named constraints, and they expire when those constraints change. A higher completion rate on a small pack does not prove general coding ability, a lower price, or a permanent free grant. The card should state the negative claims in plain language, including what was not measured and which arms were absent. Publishing the hash next to the table lets a later reader see whether the pack moved after the headline was written.

Failure analysis belongs in the same narrative as the score, because a green fraction without its misses invites overreading. A timeout cluster may say more about queue behavior on a free server than about the agent's repair skill. A pass that appears only after a second retry tells a different story from a pass on the first frozen attempt. Recording those patterns beside the median keeps the result closer to a lab note than to a launch sentence.

Synthetic trials can teach the reader how to hold that line before a live route is involved at all. Suppose two repeats of one holdout item pass on the first seed and fail on the second, while wall time jumps from a quiet run to a queued run. The completion field may still look acceptable if only the last flag is counted, but the pass spread exposes the instability. A reviewer who sees that spread should keep the lane label attached and decline to promote the fraction into a ranking table.

Who should leave this method alone

This approach should not be used to certify a model for production, to claim a security review, or to rank vendors for a purchase committee. It also fails when the task pack is private, when tests depend on live paid APIs, or when the author edits fixtures after seeing failures. A free server option can disappear, slow down, or change routing, so a result from one afternoon is a sample of that afternoon. Teams that need a contractual service level should repeat the same card on an explicit environment before they trust the delta.

An engineer rehearsing this card can check MonkeyCode's current free model access and free server option before sharing a score. The live terms should be copied into the lane label before any score is shared with another reader. The useful habit is the refusal to rank an unlocked pack, and that habit remains valuable even if the hosting choice changes. A later reader should be able to reject the number without guessing which track, which wind, or which class of car produced it.

Top comments (0)