DEV Community

Jordan Liu
Jordan Liu

Posted on

I Froze the Hash. The Free Run Had to Earn the Sign.

A free runner is not a result. A stopwatch is not a race, and a pass rate is not a decision until the sample, the scorer, and the stop rule are already boring. If any of those can still wriggle, a free model and a free server only make the wriggling cheaper.

Freeze the hash. Pre-register the break. Let a sign appear only when both halves agree.

You know the itch. A stack shows up, the calls cost nothing you can feel, a server is offered so you skip provisioning, and lunch starts to look like a leaderboard. Did you ever nudge one prompt, rerun, and then describe the last number as if it had been the hypothesis?

If that scene feels familiar, it is the habit this card is built to interrupt. Not a trophy. A failure mode with a friendly face.

The feed keeps congratulating small models for going quiet on a private checker. Quiet is not stable. A checker can go silent because you edited the items until it did. Have you ever watched a green cell calm a room that never asked whether the other half of the file would have stayed green?

Seasoning and judging are different jobs. Taste the soup to change the recipe, and you are cooking. Taste it to award a ribbon, and you are peeking. Free access is a gift to the cook and a trap for the judge, because another taste barely stings.

The chart gets patient. You do not.

Disclosure: This article was prepared as part of MonkeyCode's product outreach. For this method I am treating two availability claims as given, and only those: free model access, and a free server option. I am not stating a token quota, a model roster, a machine size, a duration, or a score from a run.

I do not have a primary source I am willing to pin next to a task hash, and I will not fill the gap with a figure that might already be stale. If it is not in the card, it is not a result. A number that cannot sit in the same commit as the tasks is just weather.

What are those surfaces doing in the method? They are a pair of runners, not a mascot. Treatment A is free model access: your process sends a frozen item and scores the completion locally. Treatment B is the free server: the same item goes to a hosted runner you do not administer, and you score the trace that comes back with the same function.

Same items, same rubric, different place the work happened. If you cannot name which surface tripped a rule you wrote yesterday, you do not have an evaluation. You have a demo with extra steps.

I have not executed this harness against that product for this article. The blocks below are a proposal with a synthetic fixture so you can see the refusal path on your laptop. Thresholds are choices you lock before the first network call. They are not measurements, and they are not a review of anyone's uptime.

The unit is the item, not the afternoon. You export tasks to JSONL, one object per line, each with an id and an expected string. You hash the file bytes, then split with a seeded digest so half 0 and half 1 stay put on rerun. You score both treatments on both halves.

A sign is earned only when each half clears the item floor, each arm stays under the transport-error ceiling, the absolute gap clears a floor you wrote down so a one-item wobble cannot wear a crown, and the two halves agree on the direction. A hash mismatch is an invalidation even if the new file is basically the same. Editing the rubric after a sad cell is an invalidation too. Git will remember, even if your recap does not.

Why not one mean? A single average can sand a flip off the table until the room feels calm. I would not trust that sanding in a write-up. Would you change a routing file because forty items leaned left and forty leaned right?

I would call that a break and leave the config alone. Agreement is a low bar. It is still a bar.

Performs well cannot mean a percentage I did not measure. Here it means the run can finish under rules another person can replay: both surfaces return scoreable traces, invalidation stays quiet, and the sign matches across halves. Breaks has to stay a specific sentence, not a mood.

A timeout is not a foolish model. A rubric that never looks at tool traces is not a broken server. A moved hash is not a product defect. It is you, still cooking.

Fold those into one red badge and the free tier takes the blame for your notebook. Three sentences and zero winners is a better log. Can you say which sentence fired? If you cannot, you are not done.

Who should skip this? Anyone who needs a flattering percentage before stand-up, because the card is built to withhold it. Anyone holding private customer text should not ship it to a free server because the meter feels kind. Anyone who needs a medical, legal, or safety judgment should not pretend a frozen string check is that judgment.

If you cannot commit the task file, the card, and the scorer in one revision, you cannot audit the claim later. Do not make the claim.

Limitations, with the soft landing removed. Half-agreement is not proof that one surface is wiser, cheaper in a way that lasts, or still offered next month. It is a stability check on a small decision. A hosted runner can insert retries, clocks, and tool hops your local call never saw, and this protocol notices only if the scorer reads those fields on purpose.

String equality is a blunt rubric. It will call a harmless paraphrase a miss, and it will miss a fluent wrong answer that happens to match. A pilot that survives the card may justify a larger pre-registered run. It does not justify a testimonial, and I am not writing one.

The workflow is short enough to do before coffee gets cold, and strict enough to ruin a dishonest chart. Write the card in the same file as the scorer, hash the JSONL, and commit both plus the hash before any call. Implement two clients that take an item and return text, and keep secrets out of the repo. Leave a dry implementation that raises, so a curious teammate cannot just see what happens against the network by running the module.

Replace the dry clients only after that commit exists. Run the module and read earned. If it is false, the result is the breaks list. Sit with it. That list is the finding.

Save this as split_half_sign.py. It is a proposal. --fixture never leaves the process, and the default clients raise so a real file cannot accidentally become a leaderboard.

python3 split_half_sign.py --fixture
sha256sum tasks.jsonl
python3 split_half_sign.py tasks.jsonl "$(sha256sum tasks.jsonl | cut -d' ' -f1)"
Enter fullscreen mode Exit fullscreen mode

That second command, left as written, should fail the transport rule. Good. That is the dry default doing its job. A green earned before you replace dry_complete would mean the harness is lying, and you should delete it.

#!/usr/bin/env python3
# Proposal, not a product benchmark.
# Lock CARD before any network call. Thresholds are design choices.
# --fixture is synthetic and opens no socket.
from __future__ import annotations

import hashlib
import json
import sys
from dataclasses import asdict, dataclass
from pathlib import Path
from typing import Callable

CARD = {
    "protocol": "split-half-sign-v1",
    "min_items_per_half": 30,
    "min_scored_fraction": 0.95,
    "max_transport_error_rate": 0.05,
    "min_abs_gap": 0.02,
    "seed": "freeze-me-before-running",
}

CompleteFn = Callable[[dict], str]


def file_hash(path: Path) -> str:
    return hashlib.sha256(path.read_bytes()).hexdigest()


def load_jsonl(path: Path) -> list[dict]:
    rows = []
    for line in path.read_text(encoding="utf-8").splitlines():
        if line.strip():
            rows.append(json.loads(line))
    return rows


def half_of(item_id: str, seed: str) -> int:
    digest = hashlib.sha256(f"{seed}:{item_id}".encode()).hexdigest()
    return int(digest[-1], 16) % 2


def rubric(task: dict, completion: str) -> bool:
    expected = str(task.get("expected", ""))
    return completion.strip() == expected.strip()


@dataclass
class ArmReport:
    name: str
    n: int
    scored: int
    transport_errors: int
    pass_rate: float | None
    invalidated: str | None


def run_arm(name: str, items: list[dict], complete: CompleteFn) -> ArmReport:
    errors = scored = hits = 0
    for task in items:
        try:
            text = complete(task)
        except RuntimeError:
            errors += 1
            continue
        scored += 1
        hits += int(rubric(task, text))
    n = len(items)
    err_rate = (errors / n) if n else 1.0
    frac = (scored / n) if n else 0.0
    invalidated = None
    if n < CARD["min_items_per_half"]:
        invalidated = "below_min_items"
    elif err_rate > CARD["max_transport_error_rate"]:
        invalidated = "transport"
    elif frac < CARD["min_scored_fraction"]:
        invalidated = "unscored_mass"
    rate = (hits / scored) if scored else None
    return ArmReport(name, n, scored, errors, rate, invalidated)


def sign_of(left: float | None, right: float | None) -> int:
    if left is None or right is None:
        return 0
    gap = left - right
    if abs(gap) < CARD["min_abs_gap"]:
        return 0
    return 1 if gap > 0 else -1


def evaluate(rows, arms, got_hash, expected_hash):
    breaks = []
    if expected_hash is not None and got_hash != expected_hash:
        return {"earned": False, "breaks": ["hash_mismatch"], "hash": got_hash, "sign": 0}
    halves = {0: [], 1: []}
    for row in rows:
        halves[half_of(str(row["id"]), CARD["seed"])].append(row)
    half_signs = []
    reports = []
    for h, subset in halves.items():
        rates = {}
        for name, fn in arms.items():
            report = run_arm(f"{name}.h{h}", subset, fn)
            reports.append(asdict(report))
            if report.invalidated:
                breaks.append(f"{name}.h{h}:{report.invalidated}")
            rates[name] = report.pass_rate
        half_signs.append(sign_of(rates.get("free_model"), rates.get("free_server")))
    if not breaks and (not half_signs or 0 in half_signs):
        breaks.append("gap_below_floor_or_unscored")
    if not breaks and half_signs[0] != half_signs[1]:
        breaks.append("halves_disagree")
    earned = not breaks
    return {
        "earned": earned,
        "sign": half_signs[0] if earned else 0,
        "half_signs": half_signs,
        "breaks": breaks,
        "hash": got_hash,
        "reports": reports,
        "note": "+1 means free_model ahead on pass rate; -1 means free_server ahead",
    }


def dry_complete(_task: dict) -> str:
    raise RuntimeError("dry run: replace me after the hash is committed")


class FixtureComplete:
    # Synthetic. Pass rates flip by half. Not a vendor measurement.

    def __init__(self, name: str) -> None:
        self.name = name

    def __call__(self, task: dict) -> str:
        h = half_of(str(task["id"]), CARD["seed"])
        want_hit = (self.name == "free_model" and h == 0) or (
            self.name == "free_server" and h == 1
        )
        return str(task["expected"]) if want_hit else "MISS"


def synthetic_rows() -> list[dict]:
    buckets = {0: [], 1: []}
    i = 0
    while len(buckets[0]) < 40 or len(buckets[1]) < 40:
        item_id = f"t{i:04d}"
        h = half_of(item_id, CARD["seed"])
        if len(buckets[h]) < 40:
            buckets[h].append({"id": item_id, "expected": f"ok-{i}"})
        i += 1
    return buckets[0] + buckets[1]


def main(argv: list[str]) -> int:
    if len(argv) >= 2 and argv[1] == "--fixture":
        rows = synthetic_rows()
        blob = "\n".join(json.dumps(r, separators=(",", ":")) for r in rows).encode()
        digest = hashlib.sha256(blob).hexdigest()
        out = evaluate(
            rows,
            {
                "free_model": FixtureComplete("free_model"),
                "free_server": FixtureComplete("free_server"),
            },
            digest,
            digest,
        )
        out["fixture"] = "synthetic_refusal_not_a_product_run"
        print(json.dumps(out, indent=2))
        return 0
    if len(argv) != 3:
        print(
            "usage: split_half_sign.py --fixture | split_half_sign.py tasks.jsonl HASH",
            file=sys.stderr,
        )
        return 2
    path = Path(argv[1])
    out = evaluate(
        load_jsonl(path),
        {"free_model": dry_complete, "free_server": dry_complete},
        file_hash(path),
        argv[2],
    )
    print(json.dumps(out, indent=2))
    return 0


if __name__ == "__main__":
    raise SystemExit(main(sys.argv))
Enter fullscreen mode Exit fullscreen mode

A task file is just lines like these. Hash the bytes you will actually send, not a prettier export you make later.

{"id":"t0001","expected":"pong"}
{"id":"t0002","expected":"pong"}
Enter fullscreen mode Exit fullscreen mode

When the fixture prints, look at three fields and ignore the urge to screenshot a winner. earned should be false. breaks should name the disagreement, and half_signs should be a pair that does not match. If a future real run prints earned true, the sign is a stable direction on this frozen file, under this blunt rubric, on this day.

It is still not a quota, not a hardware review, and not permission to skip the private-data rule. Puppet rates inside --fixture are not a vendor score. They exist so you can watch the harness keep its mouth shut. When you later point the same function at real clients, silence still means silence.

I will ask for one thing, then stop selling. If free model access and a free server are actually available on the MonkeyCode account you use, point this card at tasks you can stand to hash in public. Commit the hash and the stop rule before the first call.

If the sign does not earn itself, write about the break. I trust that note more than a last-run percentage.

Top comments (1)

Some comments may only be visible to logged-in visitors. Sign in to view all comments.