DEV Community

Jordan Liu
Jordan Liu

Posted on

I Held One Lever Still

A free model and a free server are two levers. Move both, then quote one pass rate, and you did not review a model. You published a blend. I will not publish a blend.

You know the afternoon this happens. The endpoint answers. The box is up. The suite is already on disk, so you fire it, crop the fraction, and feel done. Which lever moved? If the honest answer is both, plus a retry you forgot to count, the fraction is a story about a pile.

Free infrastructure makes the pile feel scientific. No invoice, no guilt, and somehow no design. I do not buy that trade. A zero bill is not a baseline. It is a quiet way to change two things and then narrate one of them.

Disclosure: This article was prepared as part of MonkeyCode's product outreach. Two availability claims were supplied for this draft, and they are the only product facts I will stand on: free model access, and a free server option. I am not printing a token grant, a model roster, a hardware sheet, a duration, or a license badge. None of those arrived as a primary source I can defend for 2026-10-10. A number I cannot cite does not become true because a draft would read warmer with a big integer in the first screen.

So what experiment is left, if the headline is banned?

A two-by-two, blocked on the same task list, with one lever held still. Same prompts. Same grader. Same seed rule. Same timeout. The only legal swap is the factor you wrote into the manifest before the first request left the machine. Did the salt change, or the pot? If you changed both and then swore the broth was brighter, you cooked an anecdote. Hold the pot. Change the salt. On the next batch, hold the salt and change the pot. Now a later reader has something to argue with.

Hold the server, swap the model slot. That cell may talk about free model access, and only against the server you actually held. Hold the model, swap the server slot. That cell may talk about the free server, and only against the model you actually held. Swap both and you have an interaction check, not a verdict. If the single swaps were calm and the pair fell over, you learned they misbehave as a couple. You did not earn "the model failed." You also did not earn "the server failed."

The baseline cell is the ruler. It can be a recorded fixture plus a local stub, as long as you can hold that pair still next week. A ruler that moves is not a ruler. It is a second treatment pretending to be the wall. I would rather compare against a boring fixture than against whatever the router felt like serving at lunch.

Blocking is the part people skip because it looks fussy. It is not fussy. If cell A sees the tasks in one order and cell B sees a shuffle, order became a third lever. If one cell retries and the other does not, retry policy became a fourth. If the seed is "whatever the process clock felt like," you cannot tell a flaky item from a factor. Derive the seed from the task id plus one run salt, write that salt into the manifest, and refuse to resample until the rate looks friendly. Friendliness is not a stopping rule. It is a wish.

What does a free path break, even when you are trying to be good? A label that is really a router. A server that sleeps, so the first tasks measure wakeup and the rest measure the work. A client that retries while the manifest still says one attempt. An SDK that returns a friendly alias, so you never see which level you hit. If the factor name is an alias, the cell is unlabeled. An unlabeled cell is a missing cell. Would you sign a lab book that said "the blue one" and meant three different blues? I would not.

I am not going to invent the quota that punches a hole in the run. I am going to treat an unfinished cell as unfinished. Survivors do not get promoted to "the run." That promotion is how a partial afternoon becomes a chart with a calm typeface. Missing cells do not get filled with the mean of their neighbors. Neighbors are different treatments. Averaging across them is how a blend learns to wear a single name.

Read four symbolic outcomes, not as data I collected, but as the only sentences the grid licenses. Suppose model-only moves and server-only does not. You may talk about the model slot. Suppose server-only moves and model-only does not. You may talk about the server slot. Suppose neither single swap moves, and the pair does. You may talk about interaction, then stop. Suppose every cell moves the same way. You probably changed the harness, the tasks, or the grader, and the levers are innocent until you prove otherwise. None of those sentences need a fake percentage. The shape is the finding. A percentage without the shape is decoration.

The artifact is a protocol, not a lab log. I have not executed it against a live free tier for this article. The client body is a stub on purpose. If you run it after you replace the stub, the JSON manifest is the result. If you do not run it, you do not get to borrow certainty from the shape of the code. A green fraction from an unimplemented client would be a magic trick. I am not collecting those.

#!/usr/bin/env python3
"""Two-by-two eval manifest. Reference protocol, not a scored bake-off."""
from __future__ import annotations

import hashlib
import json
import os
import time
import uuid
from dataclasses import asdict, dataclass, field
from typing import Callable, Optional

@dataclass
class Factor:
    name: str
    level: str
    endpoint: str

@dataclass
class Trial:
    task_id: str
    seed: int
    model_level: str
    server_level: str
    ok: Optional[bool] = None
    error_class: Optional[str] = None
    latency_ms: Optional[int] = None

@dataclass
class Manifest:
    run_id: str
    harness_sha256: str
    task_file_sha256: str
    seed_salt: int
    started_unix: float
    cells: dict = field(default_factory=dict)
    finished_unix: Optional[float] = None
    admissible: bool = False
    refusal_reason: Optional[str] = None

def sha256_file(path: str) -> str:
    return hashlib.sha256(open(path, "rb").read()).hexdigest()

def classify(exc: BaseException) -> str:
    if isinstance(exc, NotImplementedError):
        return "unimplemented"
    name = type(exc).__name__
    if "Timeout" in name:
        return "timeout"
    if "Connection" in name or "HTTP" in name:
        return "transport"
    return "grader_or_task"

def run_cell(tasks, model: Factor, server: Factor, salt: int, call: Callable) -> list[Trial]:
    out = []
    for task_id in tasks:
        seed = int(hashlib.sha256(f"{salt}:{task_id}".encode()).hexdigest()[:8], 16)
        trial = Trial(task_id, seed, model.level, server.level)
        t0 = time.perf_counter()
        try:
            call(task_id, model, server, seed)
            trial.ok = True
        except Exception as exc:
            trial.ok = False
            trial.error_class = classify(exc)
        trial.latency_ms = int((time.perf_counter() - t0) * 1000)
        out.append(trial)
    return out

def judge(manifest: Manifest) -> tuple[bool, str]:
    need = {"baseline", "model_only", "server_only", "both_swapped"}
    have = set(manifest.cells)
    if have != need:
        return False, "missing_cells:" + ",".join(sorted(need - have))
    for name, trials in manifest.cells.items():
        if not trials:
            return False, f"empty:{name}"
        if any(t["error_class"] == "unimplemented" for t in trials):
            return False, "stub_call_is_not_evidence"
        if any(t["ok"] is None for t in trials):
            return False, f"unscored:{name}"
    return True, "grid_complete"

def rates_if_admissible(manifest: Manifest) -> Optional[dict]:
    ok, _reason = judge(manifest)
    if not ok:
        return None
    report = {}
    for name, trials in manifest.cells.items():
        graded = [t for t in trials if t["error_class"] is None]
        report[name] = {
            "n": len(trials),
            "graded_n": len(graded),
            "pass_rate": (sum(bool(t["ok"]) for t in graded) / len(graded)) if graded else None,
            "transport_n": sum(t["error_class"] == "transport" for t in trials),
            "timeout_n": sum(t["error_class"] == "timeout" for t in trials),
        }
    return report

def main() -> None:
    model_base = Factor("model", "baseline-slot", os.environ["MODEL_BASE_URL"])
    model_free = Factor("model", "free-model-slot", os.environ["MODEL_FREE_URL"])
    server_base = Factor("server", "baseline-slot", os.environ["SERVER_BASE_URL"])
    server_free = Factor("server", "free-server-slot", os.environ["SERVER_FREE_URL"])
    task_path = os.environ["TASK_FILE"]
    tasks = json.load(open(task_path))
    salt = int(os.environ.get("EVAL_SEED", "17"))

    def call(task_id: str, model: Factor, server: Factor, seed: int) -> None:
        raise NotImplementedError(
            f"wire your client: {task_id} {model.level} @ {server.level} seed={seed}"
        )

    plan = {
        "baseline": (model_base, server_base),
        "model_only": (model_free, server_base),
        "server_only": (model_base, server_free),
        "both_swapped": (model_free, server_free),
    }
    manifest = Manifest(
        str(uuid.uuid4()), sha256_file(__file__), sha256_file(task_path), salt, time.time()
    )
    for name, (model, server) in plan.items():
        manifest.cells[name] = [asdict(t) for t in run_cell(tasks, model, server, salt, call)]
    ok, reason = judge(manifest)
    manifest.admissible = ok
    manifest.refusal_reason = None if ok else reason
    manifest.finished_unix = time.time()
    json.dump(asdict(manifest), open("run_manifest.json", "w"), indent=2)
    if not ok:
        raise SystemExit(f"refusing headline metric: {reason}")

if __name__ == "__main__":
    main()
Enter fullscreen mode Exit fullscreen mode

Pin the harness and the task file before you trust the run you think you started. A dirty tree is another lever, and it never introduces itself.

sha256sum eval_grid.py tasks.json | tee inputs.sha256
export MODEL_BASE_URL MODEL_FREE_URL SERVER_BASE_URL SERVER_FREE_URL
export TASK_FILE=tasks.json EVAL_SEED=17
python3 eval_grid.py
python3 -c 'import json; m=json.load(open("run_manifest.json")); print(m["admissible"], m.get("refusal_reason"))'
Enter fullscreen mode Exit fullscreen mode

A stub run should exit on stub_call_is_not_evidence. That refusal is the correct first result. Do not catch it and tweet the partial JSON. The partial JSON is a note that the client was never wired.

When a real client is in place, debug the manifest before you debug the model. Open run_manifest.json and compare transport_n across cells. If transport piles up only in server_only and both_swapped, your sentence is about the server path, or about the path to that server. It is not about the model. If timeouts pile up only on the free model slot, say timeout, then decide whether your product even cares. If missing_cells fires, you do not have a result. You have a plan you did not finish. Imputing the missing corner from the other three is how people invent an interaction they never observed.

What do you commit? The harness, the task-file hash, the seed salt, and the manifest. Not a screenshot of a fraction. A screenshot cannot tell me whether the both-swapped cell was allowed to speak. The JSON can. If admissible is false, the writeup is the refusal reason, and that writeup is allowed to be short. Short and refused beats long and blended.

Who should walk past this? You, if the demo is at lunch and the room wants a vibe. A grid is a bad vibe. You, if the question is whether a UI feels kind, because kindness is not a factor level in this file. You, if every request is locked to one blended endpoint and you cannot hold a baseline still. Then you have one pipe. A pipe can still be worth using for a prototype. It cannot carry a sentence that starts with "the model," and it cannot carry one that starts with "the server." Say "this pipe" and stop. That sentence is weaker, and it is the one you actually earned.

Limits, with the romance removed. There is no scoreboard here, because I did not run the live pair, and I will not invent one to decorate a method. Grader quality sits outside the grid. A stable wrong grader produces an admissible, repeatable, useless manifest. Environment variables are not a docs site. Free access can move next week, so last week's JSON does not license next week's slogan. Interaction can be real and still be too small to matter for your task. The protocol will not choose your practical threshold. Write that threshold down before you look, or do not pretend the look was a test.

Trying the pipe is fine. Publishing the try as a model review is the part I will not cosign. If the free model slot and the free server slot in your env file are the ones named in the disclosure, they are eligible factors in this grid. They are not pre-cleared results. Run the four cells once, keep the refusal or the JSON, and let that file outlive the badge.

Top comments (0)