A remote agent trace is a reliable diff only after every tool call has been closed on the machine that ran it, and only after arguments and results have been replaced by stable hashes. Exported span text is a label. The label can match while the body was truncated, reordered by a collector, or dropped to fit an attribute budget you do not control.
That constraint shows up as soon as a demo leaves a laptop. Public developer discussion this week has been heavy with agent demos, hackathon calls, and portfolios assembled quickly with model help. Those pieces can be fine as demos. They are a poor debug log. A green summary after two runs does not tell you whether the same tool was invoked, whether the arguments matched, or whether a collector kept the fields you meant to compare.
Treat the export path like a courier that photographs the crate label and may discard the crate. If the photograph is all you archive, the next regression will argue about photography. The loop below keeps the crate in a local store and ships the photograph on purpose. A warehouse of labels can still tell you that a call named apply_patch left with status ok. It cannot tell you which bytes were in the patch if the photo cropped the middle of the label. Teams then "fix" the diff by comparing cropped strings, and the crop quietly becomes the spec.
The loop has four beats, and it stays on one machine until the last beat. First, wrap the tool boundary so entry and exit are recorded in the same process that executed the call. Second, redact, canonicalize, hash, and write the canonical body to a content-addressed directory. Third, refuse export while any call lacks a terminal status or any hash lacks a local file. Fourth, diff two runs by the multiset of tool name, status, argument hash, and result hash. A summary string is not an input to that diff. Length rides along so a 40 KB patch and a 400 byte patch cannot masquerade as the same event before anyone opens the body files.
Canonicalization is the part people skip, and it is the part that makes the hash mean something. JSON objects do not have a stable key order. A model that emits one key order and a wrapper that emits another are the same call if you sort keys and fix separators before hashing. They are a false regression if you hash the raw string a logger happened to print. Timestamps, random request ids, and absolute temp paths do not belong inside the hashed body. Strip them into attributes that you do not use for equality, or the gate will fail closed on noise.
Redaction sits before hashing, not after. If an argument contains a token, a customer record, or a signed URL, replace that field with a typed placeholder and hash the placeholder form. Hashing a secret and exporting only the hash is better than exporting the secret, but the local store would still hold the secret if you write the raw value. Write the redacted canonical form, and keep the real secret in a system that already has an access story. This note is not a redaction product. It only fixes the order: redact, canonicalize, hash, store, then consider export.
The listing that follows is a proposal. It was not executed against a live collector for this draft, and it reports no timings, quotas, or pass rates. Non-JSON payloads need a separate byte canonicalization. This version assumes JSON-compatible values.
#!/usr/bin/env python3
"""Local hash gate before an agent trace export. Proposal, not a measured run."""
import hashlib, json, pathlib
from collections import Counter
def canonicalize(value):
return json.dumps(value, sort_keys=True, separators=(",", ":"), ensure_ascii=False)
def payload_hash(value):
body = canonicalize(value).encode("utf-8")
return hashlib.sha256(body).hexdigest(), len(body)
def tool_span(name, args, result, status):
args_hash, args_len = payload_hash(args)
result_hash, result_len = payload_hash(result)
return {
"name": name,
"status": status,
"args_sha256": args_hash,
"args_len": args_len,
"result_sha256": result_hash,
"result_len": result_len,
}
def write_bodies(store, span, args, result):
root = pathlib.Path(store)
root.mkdir(parents=True, exist_ok=True)
for digest, body in ((span["args_sha256"], args), (span["result_sha256"], result)):
path = root / digest
text = canonicalize(body)
if not path.exists():
path.write_text(text, encoding="utf-8")
elif path.read_text(encoding="utf-8") != text:
raise RuntimeError(f"unstable canonical form: {digest}")
def export_ready(spans, store):
problems, root = [], pathlib.Path(store)
for span in spans:
if span["status"] not in {"ok", "error"}:
problems.append(f"open:{span['name']}")
for key in ("args_sha256", "result_sha256"):
if not (root / span[key]).is_file():
problems.append(f"missing:{span[key]}")
return problems
def manifest(spans):
return [(s["name"], s["status"], s["args_sha256"], s["result_sha256"]) for s in spans]
def diff_manifests(left, right):
cl, cr = Counter(map(tuple, left)), Counter(map(tuple, right))
return {
"only_left": sorted((cl - cr).elements()),
"only_right": sorted((cr - cl).elements()),
}
A command-shaped check, still illustrative, is short enough to alias. Record two fixtures, refuse the push when the gate prints any problem, and diff manifests only after both gates are empty.
python trace_gate.py run --fixture fixtures/rename_symbol.json --store runs/a/bodies --out runs/a/spans.json
python trace_gate.py gate --spans runs/a/spans.json --store runs/a/bodies
python trace_gate.py diff --left runs/a/spans.json --right runs/b/spans.json
The run subcommand is intentionally not fully spelled out. Wire it to the agent runner so the same function both invokes the tool and calls tool_span plus write_bodies. A logger that scrapes stdout after the fact will hash a formatted line, not the value the tool received, and the gate will bless a fiction. Fixture choice matters more than collector choice. A fixture that calls the live network will hash a moving world. Pin the tool doubles. A rename fixture that reads three files from a temp repository, applies one patch, and writes a unified diff is enough. On the next run, a model that reflows the patch shows up as only_right, which is the signal you wanted, not a collector glitch.
What leaves the machine is the span record: tool name, status, two hashes, two lengths. What stays is the body directory. That split is the decision rule. If a body is small, already redacted, and you have a written reason to debug it remotely, you may attach it as an attribute. If you cannot state the collector's current attribute budget, do not attach it. A budget you guess will be wrong on the day the cap tightens, and a truncated JSON attribute will not match the local body. Identical runs then look flaky, and the flake is in the courier.
MonkeyCode's free model access and free server option, as described for this workflow by the operator, fit on the two sides of that split. Free model access is a way to repeat a pinned fixture often enough that a hash diff has more than one sample. A free server is a place the reduced spans can land when you do not want to stand up a collector just to see the manifest. Disclosure: This article was prepared as part of MonkeyCode's product outreach. This draft does not restate model names, token quotas, hardware, retention, or duration. Those figures change, and an article is the wrong cache for them. Read the current project documentation before you depend on a limit, and treat anything not written there as unknown.
The free server does not relax the gate. A hosted collector is a stronger reason to hash locally, because you cannot assume which attributes survive the hop, and you should not put raw tool arguments on a machine you do not operate just to make a diff easier. Replay still happens from the local body store. The remote copy is an index. After the server accepts a batch, fetch the stored spans once and check that each args_sha256 you sent is still present and unchanged. If the read-back drops a hash, stop using that path for regression. A missing hash is not a partial success. It means the courier photographed a label and then smudged it. Ordering across hosts is a separate check. Do not fold clock skew into the hash, or every machine becomes a unique run.
Hash equality will not save a semantic bug. Two patches can differ by a comment and produce different result hashes while both fix the build. Two identical hashes can still hide a bad patch if the fixture never executed the failing path. The gate answers a narrow question: did the tool boundary see the same recorded inputs and emit the same recorded outputs as the run you kept? It does not answer whether the outputs were wise. Generation text that never passed through a tool needs its own fingerprint, stored beside the manifest, not stuffed into the tool hash. When the manifest diff is empty and the read-back matches, you still have not proved the agent is fixed. You have proved the tool boundary was stable across the two runs you kept. Hold the loop to that smaller bar.
Do not use this loop where the arguments are data you are not allowed to write to local disk either. A content-addressed store is still a store. Do not use it as a compliance archive. Nothing here states a retention period, an access policy, or a deletion story for a free server. Do not use it to judge sampler-only bugs. If the only difference between runs is a temperature draw inside the model, the manifests will flap and the flap will look like a tool regression. Fix the fixture, or stop diffing those fields. And do not bypass the gate because a demo deadline wants spans on a dashboard tonight. That bypass is how truncated attributes become the source of truth.
The useful close is boring, which is what you want from a debug loop. Keep the bodies next to the run, export the hashes, and let the remote index earn trust only after a read-back. If that index is a free server, confirm today's limits in the docs, then point the exporter at the reduced record rather than at the crate.
Top comments (0)