A free call that comes back with HTTP 200 is not a model that can do the job. I keep repeating that because a single pass rate keeps crawling back into the notes, wearing a lab coat. Reach, contract, and check are three different arguments. Glue them into one boolean and you will praise a host for being awake, or blame a model for a timeout it never saw.
You know the Tuesday demo. Same prompt, Wednesday, and the body is empty, or half an object, or a charming paragraph that fails the checker you trusted last month. Which piece actually broke? If you stored one bit, you cannot say. You graded a smoothie and called it fruit.
This is a protocol, not a scoreboard from a run I am about to invent. No latency trophy. No win rate. Those numbers belong to a log you produce yourself. What I will hand you is the split, a harness, and a rule for which column is allowed to speak.
The blend is the bug
Picture a clinic that publishes one word, healthy. Blood pressure, whether the patient showed up, and whether the vial was labeled all collapse into that word. A missed appointment looks like a bad heart. A mislabeled vial looks like a no-show. You would fire that clinic. Why do we publish model evals that way?
Free model access makes the blend worse, not kinder. A free server is a second moving part. It can refuse, queue, reset, or idle-timeout while the weights sit there innocent. A free model can also be perfectly reachable and still emit prose where your tool wanted an object. Score only the final checker and both failures wear the same red mark. Then the postmortem becomes a vibe.
I am not stating a quota, a model list, a box size, or a promise that any free tier stays free next week. Those facts rot on contact. If a vendor page is your source, read it the day you run, and paste that URL into the run record. Do not borrow a sentence from a blog as a spec.
Disclosure: This article was prepared as part of MonkeyCode's product outreach. The operator supplied two availability claims I will use as conditions, not as trophies: free model access, and a free server option. I was not given a verified catalog, a token ceiling, or a hardware note, so this draft does not invent them. MonkeyCode is one place those two conditions might be obtained. The protocol does not depend on that product. Swap the transport and the lesson still stands.
Freeze the exam before you touch the wire
I want one lever moving. The prompt set is a directory of files. Hash each file, then hash the sorted list of those digests. That second digest is the exam id. If a prompt changes after the first request, you are holding a different exam, and the comparison is theater. Throw the run out. Yes, that is fussy. It is also the only way a later diff means anything.
The checker is frozen too. A second function, local, pure, no network. It reads the parsed output and the fixture, then returns pass or fail plus a short reason. If the checker calls a model, you built a mirror. Mirrors always agree with themselves. I will not score that and call it evidence.
Conditions are a pair, not a tour of every host you can find. Condition A is the control: same exam id, a transport you already trust, even if that transport is a process on your laptop. Condition B is the free-model call landing on the free server. Same items. Same checker. Same timeout, written in milliseconds before the first request leaves. Edit the timeout after you see the red rows, and you edited the question.
I do not retry inside this score. A retry is a different policy, with its own article and its own lies. Score the first completed attempt. Log the transport error beside it so a later retry study can reuse the raw row without contaminating this one.
Three columns, and the order is the point
Reach asks a rude question. Did the transport return a body you are allowed to grade? DNS failure, connection reset, non-200, empty string, cutoff by the timeout you pre-registered: those are reach failures. They are not model failures. Write them in the reach column and stop. Do not feed a timeout string into the checker and call the item wrong. That move is how a flaky path becomes a fake intelligence regression.
Contract asks a picky question. Given a body, does it match the shape you demanded? I use a JSON schema on tool-shaped tasks. Missing keys, wrong types, a paragraph wrapped around an object, a brace cut off at the end: contract failures. The model can sound sharp and still fail this. Your agent will not care that the prose was warm.
Check asks the only question a user actually meant. The frozen checker says the content is acceptable. A schema-valid lie fails here. A schema-invalid truth fails earlier, in contract, and never reaches check. That ordering is the whole method. You stop operating on the wrong organ.
A run is comparable only when both conditions share the exam id, the checker digest, the timeout, and the item ids. Miss one of those and you have an anecdote. Anecdotes are fine at lunch. They are not a release note.
A harness you can run, with no fake log under it
The code below is a proposal. I did not execute it for this draft, and I will not paste imaginary rates under it. Fill the transport with a client you really have. Point the free side at a free server only after you have confirmed, from a current primary page, that the option is still offered.
mkdir -p eval/items eval/runs
python3 -m pip install --user jsonschema
# one JSON fixture per item: prompt, schema, fixture
python3 split_score.py --items eval/items --out eval/runs/$(date -u +%Y%m%dT%H%M%SZ)
#!/usr/bin/env python3
"""Proposal harness. Unexecuted here. Ships with no sample metrics."""
import argparse, hashlib, json, time
from pathlib import Path
TIMEOUT_MS = 20000 # pre-register; do not edit after red rows appear
def digest(blob: bytes) -> str:
return hashlib.sha256(blob).hexdigest()
def load_items(folder: Path):
items = []
for path in sorted(folder.glob("*.json")):
raw = path.read_bytes()
item = json.loads(raw)
item["_id"] = path.stem
item["_digest"] = digest(raw)
items.append(item)
joined = "\n".join(i["_digest"] for i in items).encode()
return items, digest(joined)
def reach_call(transport, item, timeout_ms):
started = time.perf_counter()
try:
body, status = transport.complete(item["prompt"], timeout_ms / 1000)
except Exception as exc:
return {"ok": False, "reason": type(exc).__name__, "body": None,
"status": None, "ms": int((time.perf_counter() - started) * 1000)}
ok = status == 200 and isinstance(body, str) and body.strip() != ""
return {"ok": ok, "reason": "ok" if ok else "empty_or_status", "body": body,
"status": status, "ms": int((time.perf_counter() - started) * 1000)}
def contract_ok(body, schema):
import jsonschema
try:
parsed = json.loads(body)
except json.JSONDecodeError:
return False, "not_json", None
try:
jsonschema.validate(parsed, schema)
except jsonschema.ValidationError:
return False, "schema", parsed
return True, "ok", parsed
def score_condition(name, transport, items, checker):
rows = []
for item in items:
reach = reach_call(transport, item, TIMEOUT_MS)
row = {"item": item["_id"], "condition": name, "reach": reach["ok"],
"reach_reason": reach["reason"], "ms": reach["ms"],
"contract": None, "check": None}
if not reach["ok"]:
rows.append(row)
continue
good, why, parsed = contract_ok(reach["body"], item["schema"])
row["contract"] = good
row["contract_reason"] = why
if not good:
row["raw_head"] = (reach["body"] or "")[:240]
rows.append(row)
continue
passed, why = checker(parsed, item["fixture"]) # local, pure, no model
row["check"] = bool(passed)
row["check_reason"] = why
rows.append(row)
return rows
def rates(rows):
n = len(rows) or 1
reached = [r for r in rows if r["reach"]]
contracted = [r for r in reached if r["contract"]]
checked = [r for r in contracted if r["check"]]
return {
"n": len(rows),
"reach": len(reached) / n,
"contract_given_reach": (len(contracted) / len(reached)) if reached else None,
"check_given_contract": (len(checked) / len(contracted)) if contracted else None,
}
def main():
ap = argparse.ArgumentParser()
ap.add_argument("--items", type=Path, required=True)
ap.add_argument("--out", type=Path, required=True)
args = ap.parse_args()
items, exam = load_items(args.items)
raise SystemExit(
"Inject control and free transports, then write rates() per condition. "
f"exam={exam} items={len(items)} out={args.out}"
)
if __name__ == "__main__":
main()
Look at rates. It will not divide checks by every item and call that accuracy. Check is conditional on contract. Contract is conditional on reach. A free server that drops half the calls can still show a pretty check rate on the half that arrived. Publish only the pretty number and you are advertising the survivors. I would rather see a ugly reach column than a polished lie.
A fixture can be tiny. Here is the shape I want on disk, so the harness and the checker share one contract:
{
"prompt": "Return JSON only. Pick a status for a build that exited 0.",
"schema": {
"type": "object",
"required": ["status"],
"properties": {"status": {"enum": ["green", "red"]}},
"additionalProperties": false
},
"fixture": {"expect": "green"}
}
def checker(parsed, fixture):
ok = parsed.get("status") == fixture["expect"]
return ok, "match" if ok else "mismatch"
That checker is boring on purpose. Boring is a feature. If you cannot write a boring checker, you are not ready to split a score, and a vibe grade will flatter you faster than a free host ever could.
How to read a log you actually produced
Suppose reach falls apart on the free server while contract-given-reach stays close to the control. That is a host story, or a client story. Fix the path, the DNS, the idle timeout, the auth header. Do not prompt harder. A better sentence will not repair a reset packet. I have watched people rewrite instructions for an hour because they never logged status codes. Don't be that hour.
Suppose reach holds and contract collapses. Now you may look at the model, and at the wrapper around it. Truncation first. A tight output cap masquerades as stupidity. Then check whether a chat template stuffed a preamble in front of the object you demanded. The harness keeps raw_head for that reason. A reason code without the body is gossip.
Suppose contract holds and check collapses. Now you may talk about competence. Even then, talk about these items, this checker, this exam id. One exam is not a personality. It is not a ranking of vendors. It is a paired difference on a frozen set.
Write a minimum item count before the run. I use 30 as a floor for a smoke comparison, and I do not treat 30 as a paper. Under that floor a single weird fixture swings the rate, and you will narrate noise with a straight face. Smaller sets are allowed. Call them small in the first line. Pretending small is large is the failure mode this whole split is meant to catch.
Compare on item id, not on two averages floating in a slide. If item 14 failed only on the free path, open item 14. The mean will not tell you that the fixture held a nested list the control happened to flatten. Pair the rows. Then talk.
What I will not let this protocol claim
It will not tell you the free server is cheap enough to ship. Cost needs a bill, and I do not have yours. It will not tell you about privacy. A free host may log prompts. If a fixture contains customer text, do not send it. Synthetic items, or you have left evaluation and walked into an incident.
It will not survive a moving exam. Tweak prompts between condition A and condition B and you measured your editing. It will not fairly grade open-ended prose. The checker has to be dull. Subjective rubrics need a different design, with a judge that is not the model under test. Using the same free model to grade itself is a mirror again. Skip it.
Who should walk away? If you need an SLA, a retention promise, or a named model pinned for a regulated audit, a free option is the wrong instrument. Buy a contracted endpoint and a change window. If you only want a weekend demo, skip the harness too. A demo can be a demo. Just do not paste it into a README and call it a measurement.
Silent retries also void the reach column. If your client library retries unless you say otherwise, the row is already dirty. Turn that off, or log each attempt as its own policy. Mixing them is how bounded and unbounded behavior get the same name.
Free model access is still a gift for this kind of work, because the expensive mistake is scoring the wrong organ, not renting a larger box. A free server is useful when condition B has to be a real network path instead of a daydream. It becomes a costume the moment you let reach define quality. If you already hold a key for that pair of conditions, including through MonkeyCode after you re-read what the docs offer this week, point the transport at it and keep the three columns apart. Publish the column that moved. Your readers can survive a narrower claim. They cannot survive a blended one.
Top comments (0)