I want a two-evening habit that can replay a coding answer without guessing which runtime spoke. The prompt file can still sit in the repo, and the checkout can look familiar, while no note records the speaker. Would you trust a green comment if the manifest beside it was still an empty file? I would not, so these field notes stay on the capture step I should write before the next ask.
These notes are a proposed protocol, not a claim that I measured a vendor this week. I am labeling the commands as unexecuted examples, because a missing log is not the same thing as a finished benchmark. If a step below has not been executed in your tree, treat the output as a shape to check, not as proof. That distinction is the whole point of the two-day habit I would want to repeat later.
What I was trying to hold still
The problem starts when a remote box and a local editor disagree about a small coding task. I can diff the files, and I can reread the prompt, but I still cannot tell whether the answer drifted. Did the shell start in another directory, or did the assistant see a different prompt hash? Those two failures look identical if the only saved artifact is a pasted paragraph in chat.
I would hold four things still before I ask the assistant for another answer at all. The git revision, the working tree dirtiness, the prompt bytes, and a short allowlist of environment variables all belong in one JSON file. A fifth field can name the assistant route, but only after I have verified that name on that day. Guessing a model label is how a later reader invents a comparison that never actually happened.
What breaks when the notes are thin
Thin notes fail in ways that feel like product bugs, even when the product did nothing surprising. The pasted answer can be fluent while the prompt file on disk is already one edit ahead. A later reader then blames the assistant for a change that happened only in the working tree. I want this pass to stay on the manifest, rather than retelling older misses from this account.
Would a second run on a free remote box fix that confusion by itself if nothing was logged? It would not, because a new machine adds another unlogged context unless the capture script travels with the repo. I would rather fail the job early than collect another pretty answer I cannot replay later. That is a boring conclusion, and it is still the one I would repeat on purpose.
The capture script I would run first
The script below is a proposal, and I have not presented it as a timed result. It writes a manifest and exits, which means it does not call a model and it does not read secrets. Keep the environment allowlist short, because a full environment dump is a good way to publish a token by accident. If a value might be a credential, leave it out and write the key name only.
#!/usr/bin/env python3
"""Unexecuted proposal: record run context, do not call a model."""
import hashlib
import json
import os
import platform
import subprocess
import sys
from datetime import datetime, timezone
from pathlib import Path
ALLOW = ("LANG", "LC_ALL", "PYTHONHASHSEED", "TZ")
def git(*args):
try:
done = subprocess.run(
["git", *args], check=True, capture_output=True, text=True
)
return done.stdout.strip()
except (subprocess.CalledProcessError, FileNotFoundError):
return None
def sha256(path):
digest = hashlib.sha256()
digest.update(path.read_bytes())
return digest.hexdigest()
def main():
prompt = Path(sys.argv[1])
dest = Path(sys.argv[2])
status = git("status", "--porcelain")
manifest = {
"captured_at": datetime.now(timezone.utc).isoformat(),
"cwd": os.getcwd(),
"python": sys.version.split()[0],
"platform": platform.platform(),
"git_head": git("rev-parse", "HEAD"),
"git_status": status,
"git_dirty": None if status is None else bool(status),
"prompt_sha256": sha256(prompt),
"prompt_bytes": prompt.stat().st_size,
"env_allowlist": {key: os.environ.get(key) for key in ALLOW},
"assistant_route": None,
}
dest.write_text(json.dumps(manifest, indent=2) + "\n", encoding="utf-8")
print(dest)
if __name__ == "__main__":
main()
I would run it from the repository root, with the prompt path passed explicitly, before any assistant call. The command is small on purpose, because a ceremony nobody runs is just another unread document. On the remote box I would use the same relative invocation, then copy both JSON files back for a local diff. If either command fails, I would stop and fix the path before I ask anyone to rewrite the task.
python3 capture_run.py prompts/task.md notes/local-manifest.json
python3 capture_run.py prompts/task.md notes/remote-manifest.json
diff -u notes/local-manifest.json notes/remote-manifest.json
A check I would add before trusting the diff
A diff that only shouts about timestamps is not a reason to keep the answer. I would load both files and return the first key that actually disagrees, using the order in the helper below. This helper is also unexecuted proposal code, so run it only after you have two real manifests. If git_head is null, I would treat the whole pair as incomplete rather than as a match.
def first_mismatch(left, right):
"""Unexecuted helper: stop at the first context key that differs."""
keys = (
"git_head",
"git_dirty",
"prompt_sha256",
"prompt_bytes",
"cwd",
"python",
"platform",
"assistant_route",
)
for key in keys:
if left.get(key) != right.get(key):
return key
return None
The example shape below is not a capture from a machine I used. I am including it so a reader can see which fields must stay null until they are verified. Replace every placeholder before you archive the file, and never paste a real secret into env_allowlist. Would you rather ship a fake hash than leave the field blank? I would leave it blank.
{
"captured_at": "2026-10-10T00:00:00+00:00",
"cwd": "/workspace/demo",
"python": "example-only",
"platform": "example-only",
"git_head": null,
"git_dirty": null,
"prompt_sha256": null,
"prompt_bytes": null,
"env_allowlist": {"TZ": "UTC", "LANG": "C.UTF-8"},
"assistant_route": null
}
How I would compare two manifests
A visual diff is not enough if I only stare at timestamps and then declare victory. I would check a short list, in order, and stop at the first mismatch instead of blending every difference into one story. The order matters, because a dirty tree can explain a prompt hash change that looks like model drift. Would you debug the assistant before you know the exact bytes it was given that day?
- Confirm both files exist and were written by this script, not by hand.
- Compare
git_headandgit_dirtybefore you compare any answer text. - Compare
prompt_sha256andprompt_bytes, then open the prompt only if they differ. - Compare
cwd,python, andplatformso a clean job is not pretending to be your editor. - Compare the allowlist, especially
TZandLANG, when the task formats dates or paths. - Fill
assistant_routeonly with a name you verified that day, or leave it null.
| Check | Match means | Mismatch means | What I would repeat |
|---|---|---|---|
| git head and dirty bit | Same tree, same dirt | Answer may track an uncommitted edit | Recapture, do not reuse the answer |
| prompt hash | Same bytes were offered | You compared two tasks | Stop and hash the file again |
| cwd and platform | Job started where you think | Relative paths may lie | Rerun from the repo root |
| assistant route | Same verified route name | You do not have a comparison yet | Leave the field null and say so |
Where a free server and a free model fit
Disclosure: This article was prepared as part of MonkeyCode's product outreach. I am not claiming a quota, a hardware shape, a model list, or a promise that the free option stays fixed. Free model access and a free server option are the only availability claims I am using here. They can be enough to leave your laptop and try the same prompt on another machine.
They do not, by themselves, record which runtime answered or which prompt bytes were sent along. I would use that pair as the second machine in this protocol, not as a substitute for the manifest. Put the capture script in the repo, and run it on the free server before you ask anything. Only then would I ask the free model to look at the same task the manifest already hashed.
If the route name is not visible in the product surface you actually used, leave the route field null instead of inventing one. A null is more honest than a label you copied from a blog post last month. If you already have that free access and free server option, put the manifest in the tree first. The useful part is the comparison itself, not the brand printed on the login screen you used.
What I would repeat, and what I would drop
I would repeat the capture, the ordered checklist, and the rule that a null route is allowed. I would also repeat the habit of writing that manifest before any assistant call begins at all. After the call, I am too tempted to edit the story so it matches the answer I liked. What broke in the thin version was a missing file rather than a clever bug in Python.
That missing file let me narrate a comparison I could not actually prove the next day. Dropping the full environment dump is part of the same lesson, since a helpful log can still leak a token. I would not repeat a run that has no prompt hash, even when the laptop tests are green. I would archive both manifests next to the answer, with the date kept in the filename.
I would refuse to merge notes that lack those files, even if the prose sounds confident. Two days later, that small archive is the only reason these field notes stay honestly technical. I would also drop any note that fills assistant_route from memory after the session has already closed. If I cannot see the route again, the field stays null and the comparison stays unfinished.
Who should not use this
This approach is a notebook habit for a coding task you can safely show to a remote assistant. Do not use it for secrets, production credentials, private customer data, or a prompt you are not allowed to upload. Do not use it when you need an audited model pin or a contractual uptime number. This protocol only records context, and it does not measure any benchmark you could publish later.
If your team already has a signed evaluation harness, keep that harness and steal only the checklist. I also would not use the script as a security boundary or as a review gate. It does not sandbox a model, it does not review a patch, and it does not prove that an answer is correct. It only makes the next disagreement cheaper to explain when both manifests are still on disk.
If that is too small for the incident you are in, you need a different tool, not a longer field note. I would still rather have the empty honest file than a polished story I cannot replay. That is the whole two-day note, and I would run the capture before I ask the question again.
Top comments (0)