I opened these notes because two runs looked identical in the log and still disagreed in the answer. The prompt sentence matched on my laptop and on the remote job, so I treated both calls as one request. They were not the same request, and that lazy comparison is the mistake the next lab day should catch. Have you ever trusted a matching string in a log and skipped the bytes that actually left the process?
What the notes blamed too early
My notes from hour two still said the model was unstable, because the remote answer wrapped JSON inside a fence. Local output was bare JSON, and a string compare called that a failure even when the inner fields matched. I also saw a system line on the server that my laptop script never placed in the messages array. Why did I keep scoring the answer before I froze the exact request that produced it?
The job log showed the prompt sentence, then a pretty-printed body, and then a redacted header block beside it. Pretty printing reordered keys, and the redaction hid whether temperature was omitted or set to a default. I copied the visible prompt into a new file and called that a fixture, which only preserved the part I already trusted. A fixture that drops sampling fields is not a replay, and I learned that from the broken compare.
By the twelfth hour I still had three explanations and not a single stored request body. One note blamed temperature, another blamed a hidden system prompt, and a third blamed the remote runtime itself. All three could be true, which is why a single angry sentence in a notebook is not a diagnosis. I needed a hash I could recompute from a file, not another paragraph of mood.
Where the second machine entered
Disclosure: This article was prepared as part of MonkeyCode's product outreach. I brought MonkeyCode in because this draft's operator notes list free model access and a free server option. I wanted a second place to send one frozen body, not a scoreboard and not a hardware review. I am not naming models, quotas, hardware, duration, or a promise that the free option remains.
Those details were not supplied for this draft, so I will not invent them to sound precise. Please check the current product docs before you rely on availability, since a free option can change. The free server is only a second runtime in this note, and it is not the lesson by itself. If you delete the product name, the same steps still work on a local stub you already control.
I would not put private prompts on a shared seat just to make the field note feel complete. You can learn the drift on fixture files before you ever open a socket to a remote host. Have you already decided the model is wrong without saving the body your client actually sent? That question is the gate I missed during the first twelve hours of these notes.
What actually broke
Three separate drifts sat under one angry diff, and none of them lived inside the answer text. The laptop omitted temperature, while the server client filled a default before the call left the process. A wrapper prepended a system message only when a flag file existed on that remote host. The log writer sorted keys for humans, so an eyeball diff could not see message insertion order.
So which of those three quiet drifts would your current job logger actually print out tonight? I stopped calling the model flaky once the canonical request hash changed between the two runs. The answer diff was real, but it was downstream of a body diff I had never stored beside the note. After I pinned the body, leftover wording stayed in the notes as model behavior, not as a code regression.
That split is the part I would repeat, even when the leftover wording change itself looks boring. These are the attempts I would not repeat, written down so the next note does not start there. Each one felt productive in the moment, and each one left the outbound body off the page. If a step cannot be recomputed from a file, it is a hunch and not a result you can share.
- I compared answer strings before I stored either outbound body, and the fence made a match look like a miss.
- I copied the visible prompt line into a file and called that a fixture, which dropped every sampling field.
- I trusted a pretty-printed log, so key order and redacted headers looked like proof the bodies matched.
- I almost rewrote the prompt before asking whether the server wrapper had prepended a system line.
The artifact I would run again
This lab uses only the Python standard library, and you can run it without opening a socket. I am labeling the timeline as a worked exercise, not as a production incident with uptime or token counts. No account profile was attached to this draft, so I am not claiming a personal outage or a customer. Save the module below as request_canon.py in an empty directory before you run anything else.
import hashlib
import json
VOLATILE = {"request_id", "created_at", "trace_id"}
KEEP_NULLS = {"temperature", "top_p", "seed"}
def canonicalize(payload: dict) -> str:
cleaned = {key: value for key, value in payload.items() if key not in VOLATILE}
for key in KEEP_NULLS:
cleaned.setdefault(key, None)
messages = cleaned.get("messages") or []
cleaned["messages"] = [
{"role": item.get("role"), "content": item.get("content")}
for item in messages
]
return json.dumps(
cleaned,
sort_keys=True,
separators=(",", ":"),
ensure_ascii=False,
)
def digest(payload: dict) -> str:
raw = canonicalize(payload).encode("utf-8")
return hashlib.sha256(raw).hexdigest()[:12]
def assert_same_body(local: dict, remote: dict) -> None:
if digest(local) != digest(remote):
raise SystemExit("body drift: refuse to score answers")
The sample fixture is the bug, not a happy path you should celebrate or paste into a dashboard. Write it as pair.json beside the module, and keep secrets and private text out of the file. The remote object adds a system message and a temperature, while the local object omits both of them. Request ids differ on purpose, and the hash must ignore them or you will chase noise forever.
{
"local": {
"request_id": "local-1",
"messages": [{"role": "user", "content": "Return a one-line status."}]
},
"remote": {
"request_id": "remote-9",
"temperature": 0,
"messages": [
{"role": "system", "content": "Reply as JSON only."},
{"role": "user", "content": "Return a one-line status."}
]
}
}
Run the check from that same directory, and expect the word drift plus two different short hashes. I have not executed this sample in the drafting environment, so treat that expected line as a reading prediction. If the hashes match on your machine, only then are you allowed to discuss the answer text. If they differ, stop scoring answers and store both canonical bodies next to the field note.
python - << 'PY'
import json
from pathlib import Path
from request_canon import canonicalize, digest, assert_same_body
pair = json.loads(Path("pair.json").read_text(encoding="utf-8"))
local, remote = pair["local"], pair["remote"]
print("local ", digest(local))
print("remote", digest(remote))
print("drift" if digest(local) != digest(remote) else "same-body")
print(canonicalize(local))
print(canonicalize(remote))
try:
assert_same_body(local, remote)
except SystemExit as exc:
print(exc)
PY
What the canonical form may hide
Volatile ids may differ without making the call different, so the hash drops three name fields. Those fields are request_id, created_at, and trace_id, and I do not treat them as prompt content. Sampling fields may not differ, so a missing temperature becomes an explicit null before the dump. Message order stays in the list, because a system line at the front is a different call.
Would you want a hash that ignores the system line just to make a dashboard look green? If you later add a tool list or a response format, add that field to the kept set on purpose. A silent default inside setdefault is still a choice, and you should review it when clients change. I would rather update the fixture than teach the hash to become generous and hide drift.
A gate, not a score
Use this table before you open a prompt editor or blame the model for a messy answer. It will not tell you which answer is better, and it should not become a public leaderboard. Read the row that matches the observation, store the next artifact, and stop at the forbidden conclusion. Have you skipped a row like this because the log line already felt obvious enough to trust?
| What you observed | What to store next | What not to conclude |
|---|---|---|
| Answer text differs | Canonical request hash for both sides | The model changed its mind |
| Hashes differ | Full canonical JSON, not the log line | The server hardware is wrong |
| Hashes match, answers differ | Both raw answers plus the frozen body | Your parser or code regressed |
| Server log hides fields | Body captured at the client before send | The visible prompt is the request |
| Body needs private data | Nothing on a shared free server | Redaction in a log is enough |
A tiny test you can keep
I also keep a short assertion so a later edit cannot fix the hash by dropping messages. Paste the file below as test_canon.py beside the fixture, then run it with Python locally. It should pass on the buggy pair, because the test expects drift rather than a false match. If someone removes the system role from the canonical list, this file is supposed to fail loudly.
import json
from pathlib import Path
from request_canon import canonicalize, digest
pair = json.loads(Path("pair.json").read_text(encoding="utf-8"))
local_body = json.loads(canonicalize(pair["local"]))
remote_body = json.loads(canonicalize(pair["remote"]))
assert digest(pair["local"]) != digest(pair["remote"])
assert local_body["temperature"] is None
assert remote_body["temperature"] == 0
assert remote_body["messages"][0]["role"] == "system"
assert "request_id" not in local_body
print("canon checks passed")
What I would repeat on the next drift
Before the next drift, I would repeat the capture step even if the prompt text already looks identical. I would also repeat the hash gate before anyone on the thread starts debating answer quality. The list below is the loop I want in the note, written as actions rather than as opinions. Would I skip a step because the remote job already printed a green health line beside the log?
- Capture the outbound JSON in the client, before a proxy or a log formatter touches the bytes.
- Drop only volatile ids, and keep explicit nulls for sampling fields so a quiet omission cannot hide.
- Hash the canonical body, and refuse every answer comparison when the two short hashes diverge.
- Treat a free server as another runtime with its own wrapper flags, not as a clone of your laptop.
- Write the note as body drift or as answer drift, and never as both conclusions in the same line.
I would also keep the fixture in the repo and review it like any other test input file. A prompt sentence without the surrounding body will lie to you again sometime next week. Would you merge a test that asserted on a screenshot instead of the input you actually sent? I would not, and that is the standard I am trying to apply to these request notes.
Limits, and who should skip this
This check does not measure quality, latency, cost, or safety of any model you call. It will not catch a server that mutates the body after your client logger has already finished. It also will not help if you cannot see the outbound payload, because the hash would bless a guess. Do not send secrets, customer text, or regulated data to a free server just to finish a replay.
You should skip this approach if you need a contractual model pin, an uptime target, or a private path. You should also skip it when policy forbids third-party inference, even if a free option is listed. A lab hash will not replace an evaluation set, a human review, or a current vendor status page. If your work cannot leave the laptop, keep the fixture diff and do not invent a remote result.
If you already have a free server seat, replay one frozen body there before you rewrite the prompt. The seat is optional, and the hash is the part I would keep if that seat disappeared tomorrow. I am not asking you to switch tools, and I am not claiming this lab beat any other workflow. Start from the pair file, prove the drift on your own machine, and only then decide where to send it.
Top comments (0)