A score delta is not evidence of a prompt change when the model stamp or the server revision also moved. A nightly harness that calls a replaceable model endpoint can mix those variables without editing a single prompt line. The recorded run should refuse that comparison until the model stamp and the server revision both match the stored baseline. That refusal is the regression signal, and it should surface before any grader assertion is allowed to run.
Teams often treat a free model route as a drop-in substitute for yesterday's candidate, then read the new mean as quality. The substitution is cheap, so it happens during incident reviews, quota pauses, and routine host maintenance. A lower score can be a real prompt regression, a weaker model, or a server image that changed tokenization. Without an identity stamp, those three causes collapse into one number and leave the on-call engineer with no stack trace.
Earlier checks already pin the fixture, the grader, and the case slot before a diff is trusted. Those pins stop a moving rubric from rewriting history, but they do not name the model that produced the completion. A fixture hash can stay stable while the endpoint silently routes to a different weight set. The missing field is the model stamp, stored beside the fixture hash rather than inferred from the score.
A record that can refuse the diff
The practical fix is a three-part record that travels with every nightly run and every manual rerun. A case manifest lists golden inputs, expected envelopes, and a content hash computed over the canonical file bytes. A candidate adapter returns the completion plus the model identifier and server revision reported by the host. A quarantine step computes a score delta only when the manifest hash, model stamp, and server revision all match the baseline row.
Think of the stamp as the lot number on a reagent bottle in a wet lab. Two assays can share one protocol card and still disagree because the reagent bottle was replaced overnight. You would not publish that disagreement as a method improvement until both lot numbers match the notebook. A prompt eval deserves the same courtesy, because a free host can rotate bottles without notifying the scoreboard.
The manifest shown below is a proposal artifact, not a captured production trace from any live evaluation host. Each case carries an input the model may see and an expected envelope the grader may see only after the call returns. Keeping the gold out of the prompt is a separate control, and this note does not reopen that argument. The new requirement is that the file bytes are hashed, so a quiet edit to an expected field cannot hide inside a model swap.
{
"suite": "envelope-v3",
"cases": [
{
"id": "refund-status",
"input": {"ticket": "T-1842", "ask": "status"},
"expected": {"tool": "lookup_ticket", "args": {"id": "T-1842"}}
}
]
}
The adapter contract stays small enough that a reviewer can reimplement the same checks without adopting a new framework. A fabricated stamp would make unlike runs look comparable, so an omitted host identifier must remain unknown. Unknown is a failed quarantine rather than a silent default that pretends yesterday's model is still attached. The following snippet is unexecuted proposal code, and it only hashes local files without placing a network call.
#!/usr/bin/env python3
"""Proposal harness: refuse score diffs when stamps differ. Not a measured run."""
import hashlib
import json
import sys
from pathlib import Path
STAMP_FIELDS = ("manifest_hash", "model_stamp", "server_revision")
def file_hash(path: Path) -> str:
return hashlib.sha256(path.read_bytes()).hexdigest()
def grade_envelope(completion: dict, expected: dict) -> dict:
missing = [key for key in expected if key not in completion]
wrong = [
key
for key in expected
if key in completion and completion[key] != expected[key]
]
return {
"passed": not missing and not wrong,
"missing": missing,
"wrong": wrong,
}
def quarantine(baseline: dict, candidate: dict) -> dict:
reasons = [
field
for field in STAMP_FIELDS
if baseline.get(field) != candidate.get(field)
]
if baseline.get("passed") not in (0, 1) or candidate.get("passed") not in (0, 1):
reasons.append("passed")
if reasons:
return {"comparable": False, "reasons": reasons, "delta": None}
return {
"comparable": True,
"reasons": [],
"delta": int(candidate["passed"]) - int(baseline["passed"]),
}
def load_run(path: Path) -> dict:
record = json.loads(path.read_text())
for field in ("model_stamp", "server_revision"):
if field not in record or record[field] in ("", None):
record[field] = "unknown"
return record
def main() -> int:
if len(sys.argv) != 4:
print(
"usage: quarantine_diff.py MANIFEST BASELINE CANDIDATE",
file=sys.stderr,
)
return 64
manifest = Path(sys.argv[1])
baseline = load_run(Path(sys.argv[2]))
candidate = load_run(Path(sys.argv[3]))
candidate["manifest_hash"] = file_hash(manifest)
decision = quarantine(baseline, candidate)
print(json.dumps(decision, indent=2))
return 0 if decision["comparable"] else 2
if __name__ == "__main__":
raise SystemExit(main())
A local check uses two fixture files that differ only in the recorded model stamp and in nothing else. The command below should exit with status 2 when the quarantine fires on that cross-model pair. Status 0 would mean the three stamp fields matched and a numeric delta was printed to standard output. Neither exit status is a claim that the prompt itself improved, regressed, or stayed semantically equal.
python3 quarantine_diff.py cases/manifest.json runs/baseline.json runs/swapped_model.json
echo "exit=$?"
python3 -c "import quarantine_diff as q; d=q.quarantine({'manifest_hash':'a','model_stamp':'m1','server_revision':'r1','passed':1},{'manifest_hash':'a','model_stamp':'m2','server_revision':'r1','passed':1}); assert d['delta'] is None; print(d)"
A baseline row that survived review can look like the object below, with passed stored as an integer. Storing passed as an integer keeps the later delta a subtraction rather than a paragraph of judgment. The model stamp is an opaque string copied from the host response, not a nickname chosen by the harness. If the host later returns a different string for the same marketing name, quarantine treats that string as a new lot.
{
"manifest_hash": "recorded-sha256",
"model_stamp": "host-reported-id",
"server_revision": "image-digest",
"passed": 1
}
Suppose the baseline passed flag is 1 and the candidate passed flag is also 1, while the model stamp changes. The naive subtraction of those flags would be zero, which looks like stability to anyone watching a chart. Quarantine returns a null delta instead, because a matched score across different stamps is not evidence of stability. A displayed zero would have been the most misleading result this particular harness is able to produce.
When the candidate keeps the manifest hash but changes model_stamp, the printed decision names that field and sets delta to null. That null has to survive the dashboard, because a coercion from null to zero hides the quarantine and recreates the bug. Store the reasons list beside the null, and alert when comparable is false rather than when a score crosses a threshold. A threshold alert assumes the runs were comparable, which is the assumption this harness exists to test.
The same refusal applies when only the server revision moves and the model marketing string stays put. A new image digest can change a tokenizer, a timeout, or a tool schema while the advertised model name remains constant. The revision field exists so that class of change cannot masquerade as an edit to the prompt. Host cost is a separate axis, and a cheaper place to run the adapter does not freeze identity.
Disclosure: This article was prepared as part of MonkeyCode's product outreach. MonkeyCode is described by the operator as an open-source project that offers free model access and a free server option. Those two availability claims are operator-supplied, and they matter because the candidate adapter needs a host cheap enough for a nightly pass. Exact token allowances, model identifiers, hardware, and offer duration are omitted, because no primary document was attached to this draft.
The proposal script returns status 0 only when the three stamp fields match and a numeric delta was printed. Status 2 means the quarantine fired, which is a successful detection rather than an unexpected process crash. A wrapper should treat status 2 as a page, and it should treat a missing input file as a separate operational error. Mixing those two outcomes inside one alert rule will train the on-call rotation to ignore the page.
Silent regressions are especially likely when the score stays flat, because a flat chart does not summon a reviewer. A model swap that preserves the pass count is invisible unless the stamp comparison runs before the chart is drawn. That is why the decision object carries a reasons list even when both passed flags happen to be identical. The chart can still be drawn later, but only from rows that the quarantine has already marked comparable.
A useful nightly flow keeps four checks in a fixed order, and none of them is a slogan about model quality. Hash the manifest first, and refuse the run when that hash is absent from the baseline index. Call the candidate adapter second, and persist the host-reported model stamp and server revision even when the completion fails. Grade the stored completion third, then run the quarantine and page a human when the runs are not comparable.
Those checks matter more than the adapter language, and a shell wrapper can call the same decision after a manual replay. Write each candidate record to a new path, and promote it only after a person accepts both stamps and the delta. Automation may append run files, but automation should not bless a new lot number on its own. A person remains the only step that is allowed to replace the reference row after reading both stamps.
A content hash of raw file bytes will change when an editor rewrites newlines or reorders keys inside the manifest. That sensitivity is intentional, because a quiet formatting pass can still alter an expected field a human skips. If stable hashes across machines matter, canonicalize the JSON before hashing and store those canonical bytes. The proposal script hashes the file as saved, so different line endings will quarantine until the bytes match.
What an exact envelope does not prove
Exact envelope equality marks a fine paraphrase as a failure, and it also misses a harmful completion that still carries the expected keys. This approach is the wrong tool for open-ended writing quality, preference ranking, or any grader that needs a second model as judge. It is the wrong tool when the host cannot report a stable identifier, since every run would quarantine and the page would become noise. Teams without a frozen case manifest should build that file first, since a stamp on a moving case set still produces an unreadable delta.
Teams that already diff inside one pinned model can add the two stamp fields without rebuilding the suite. Backfill the string unknown on old rows, and treat unknown as non-comparable until a fresh baseline is recorded under observation. Do not backfill a guessed marketing name from memory, because a guessed stamp creates false permission to compare. A missing stamp is an honest block, while a guessed stamp is a quiet way to reintroduce the original bug.
If current records already store a host identifier, compare two saved rows and see whether a server move would have changed that string. One comparison is enough to tell whether the quarantine would have fired on last week's nightly job. The check does not require a new rubric, a new vendor, or any claim about which model is better. It only asks whether the score already on the chart was earned under a single lot number.
Top comments (0)