A repeated tool span with the same argument hash and a prior failed exit is not new evidence. It is the same failure, recorded twice. When the next step would spend a free model call on a free server, that repeat is the wrong unit of spend. The chat summary cannot make the second attempt distinct.
Agent debug loops still treat the final message as the thing to judge. That message is a compressed claim. The checkable unit is smaller: the tool name, a fingerprint of its arguments, the tool version, the process exit, and the diff left in the tree. If those fields already sit in a closed local row, another remote run cannot change the conclusion. It can only add latency and another unread log.
A returns desk that restocks a cracked mug because the receipt says completed has made the same mistake. The mug is the artifact. The receipt is the summary. A trace has to follow the mug. Public discussion this week leans toward portfolios, meetups, and profile showcases. Those are fine topics. They do not tell a reviewer whether a tool call already failed.
Close the span, then ask
The module below is an unexecuted proposal. It is not a measured tracer, and it is not a claim about any vendor's export format. It keeps a JSON ledger on the laptop. Raw prompts and raw arguments stay out of the file. A short SHA-256 prefix is enough to recognize a copy-paste retry inside one debug session.
#!/usr/bin/env python3
"""Unexecuted proposal: veto a retry from a local tool-span ledger."""
import hashlib
import json
from pathlib import Path
LEDGER = Path(".trace/tool-span-ledger.json")
def fingerprint(tool: str, arg_text: str, tool_version: str) -> str:
payload = f"{tool}\n{tool_version}\n{arg_text}".encode()
return hashlib.sha256(payload).hexdigest()[:16]
def load() -> dict:
if not LEDGER.exists():
return {"spans": []}
return json.loads(LEDGER.read_text())
def record(tool, arg_text, tool_version, exit_code, diff_stat, trace_id, span_id):
entry = {
"fp": fingerprint(tool, arg_text, tool_version),
"tool": tool,
"tool_version": tool_version,
"exit_code": exit_code,
"diff_stat": diff_stat,
"trace_id": trace_id,
"span_id": span_id,
}
data = load()
data["spans"].append(entry)
LEDGER.parent.mkdir(parents=True, exist_ok=True)
LEDGER.write_text(json.dumps(data, indent=2) + "\n")
return entry
def veto_retry(tool: str, arg_text: str, tool_version: str) -> str:
fp = fingerprint(tool, arg_text, tool_version)
prior = [s for s in load()["spans"] if s["fp"] == fp]
if not prior:
return "allow: no prior span"
last = prior[-1]
if last["exit_code"] != 0:
return f"block: span {last['span_id']} already exited {last['exit_code']}"
if last["diff_stat"] in ("", "0 files changed"):
return f"block: span {last['span_id']} closed with an empty diff"
return "allow: last span exited 0 with a non-empty diff"
veto_retry is the decision, not a suggestion buried in a paragraph. allow means a new remote run can add information. block means the next model call would restate a span the ledger already closed. An empty diff counts as a failed advance even when the process exit is zero. A green exit that changes no file does not move a code fix forward.
A stderr line joins that row only when it carries the same trace_id and span_id. Otherwise it remains a side channel. Unjoined logs can explain a failure after the veto. They cannot overturn the veto by themselves. A working-tree stat joins the same way: git diff --stat is stored as text on the span, not inferred later from a summary that forgot the path.
One failed formatter, then a stop
Take a formatter invoked as ruff check src at tool version 0.6.9. The child exits 1. git diff --stat prints nothing, because the tool only reported diagnostics. That pair is the evidence. A later agent turn that proposes the same command, the same version, and the same arguments is the same span in a new coat of prose.
mkdir -p .trace
python3 - <<'PY'
from span_ledger import record, veto_retry
record("ruff", "check src", "0.6.9", 1, "", "7f3a", "c21")
print(veto_retry("ruff", "check src", "0.6.9"))
print(veto_retry("ruff", "check src --fix", "0.6.9"))
PY
The expected lines are block: span c21 already exited 1 and then allow: no prior span. The second call changed the argument text, so the fingerprint changed. That is the only kind of retry this loop treats as new work. A fresh trace id is not new work. A fresh span id is not new work. Identity lives in the tool, the version, and the argument text.
The shell fragment assumes the proposal lives in span_ledger.py on PYTHONPATH. It has not been run as a published benchmark. Wiring argparse is optional. The order is not optional: record the closed span, run the veto, and only then decide whether a remote execution is eligible. Reversing that order is how a free run gets spent on a transcript the laptop already settled.
If the tree was dirty for reasons the tool did not own, add a tree hash to the fingerprint inputs before trusting allow. A dirty file that the argument text never mentions will otherwise hide inside a repeated hash. The ledger is only as honest as the fields it is given.
Spend the free run only on allow
Disclosure: This article was prepared as part of MonkeyCode's product outreach.
MonkeyCode belongs in this loop only as a destination for an eligible retry, not as the source of the verdict. The operator states that free model access and a free server option exist. This draft does not state a token quota, a hardware shape, a retention window, a model name, or a license, because those figures were not checked against a primary source on 2026-10-08. A dashboard number stays unread until it is checked on the day of the run.
The sequence is local first. Close the tool span. Write the ledger row. Run veto_retry. Point only an allow result at the free server. When the remote summary returns, join it to the new span id before anyone calls the bug fixed. If the export lacks an exit code, the local process status wins. A missing export is an incomplete span, not a success.
That order still holds if the product name is removed. Any model route with a limited run budget has the same constraint. The free option only makes the cost visible: the next person in the queue waits, or the same person spends the run that should have tested a changed fingerprint. If free model access and a free server are already on the bench, send them the allow path and leave the block path on disk.
What a 16-hex prefix can say
No timing result is claimed, and no model comparison is implied. Sixteen hexadecimal characters are 64 bits of a SHA-256 digest. That width is enough to catch copy-paste retries in a single laptop session. It is too weak to serve as a global identifier across teams. Collisions are not the decision anyway. The stored tool name and version sit beside the hash, and the veto reads all three.
The rule does not measure model quality. It measures whether the tool attempt changed. Two drafts can differ in prose and still call the same tool with the same arguments. The ledger blocks that second call on purpose. If the failure was a flake, a full disk, or a clock jump, the operator has to change an input the fingerprint can see: the arguments, the tool version, or a recorded tree hash. Until one of those changes, another sample is not a new experiment.
Secrets are the hard limit. Argument text is hashed, not stored, but the digest is still derived from a secret if a token was passed on the command line. Do not feed raw credentials into fingerprint. Redact first, hash the redacted form, and keep the redaction rule stable. A shifting redaction makes every retry look new, which quietly disables the veto.
Who should skip the ledger
Skip it when a trace backend already enforces idempotency keys on tool calls and the team pages from those keys. A second local file will drift, and drift is a second source of false certainty. Skip it for live incident response, where repeating a call may be the mitigation rather than a wasted debug retry. A veto built for code-fix loops will block a probe that was supposed to differ in network state.
Skip it when the tool output depends on time, an unread file, or a service the fingerprint cannot see. In that case the veto refuses a run that might have changed. That refusal is safer than a false fix, and it is still the wrong tool for an exploratory probe. Also skip it when the goal is a public demo that needs visible motion. This loop is a refusal mechanism. The session will look quieter. Quieter is the result.
The artifact to keep is the blocked line: span id, exit code, and the reason another free run was not started. A summary that never names that span is not a close. It is the receipt, waiting for someone to look at the mug.
Top comments (0)