DEV Community

Sam Sun
Sam Sun

Posted on

Hold the Score Until the Collector Acks the Link Set

A split-host agent run is not ready to score until the collector acknowledges every linked span you expected to export. Process exit only proves the local loop stopped. The trace you will compare is whatever a remote collector actually stored.

That gap shows up as soon as the model hop and the tool hop do not share a process. The model request is created on a service you do not parent. The tool span is created beside the working tree. A third host, the collector, knows only what arrived.

Scoring from the local log grades a draft. This note treats the gap as a hard gate. The code below is a proposed workflow, not an executed benchmark, and it reports no measured miss rate.

A parent pointer is the wrong join

Parent and child spans assume one process tree. You do not have that tree when the model call leaves the machine. Writing a local parent id onto a remote request invents a hierarchy the collector cannot check.

The relation you can defend is a link. This tool span is associated with that remote request id, and both ids must be present before a score is written. A link does not pretend the remote process was your child. It records a join key you can test.

The picture is a warehouse manifest beside a dock receipt. The manifest lists what left. The receipt lists what the dock accepted. A diff that reads only the manifest will count crates still on the truck.

A successful model status does not close the receipt. It says the model answered. It does not say the tool span, the link, and the model span landed in one trace. Wall-clock order does not close it either, because the two hosts do not share a clock you control.

Fingerprints that survive a model swap

A low-cost model swap helps a debug loop, and it ruins a naive diff. If the span name embeds a model id, two runs that called the same tools look unrelated. The label moved. The tool contract may not have moved at all.

Key the tool contract, not the vendor string. Canonicalize the tool name and the arguments, then hash them. Put the hash on the local span, and hash the result only when it is small and non-sensitive.

The comparison then asks whether the fingerprint sequence changed. It does not ask whether the model label changed. Stable hashes keep a model swap from looking like a regression.

Canonical form is part of the method. Sort keys, reject non-finite floats, and freeze separators. Otherwise a harmless reorder looks like a behavior change. Keep the canonicalizer in the same file as the gate so the expected set and the later diff cannot drift.

Raw arguments should not ride to a shared collector when they might hold secrets. The fingerprint can travel. The body stays in a local file you open only after a hash differs. That split is a limit of this loop, not a full redaction program.

The gate, as proposed code

The script has not been run for this article. It encodes one rule. Do not emit a ready score unless the ack covers every expected id, including the remote request id. Save it as ack_gate.py if you want the shell step below to import it. The type hints assume Python 3.9 or newer.

#!/usr/bin/env python3
"""Proposed gate. Unexecuted example: hold the score until the ack covers the set."""

import hashlib
import json
import sys
from pathlib import Path

def canonical(payload: dict) -> bytes:
    return json.dumps(
        payload,
        sort_keys=True,
        separators=(",", ":"),
        allow_nan=False,
    ).encode()

def fingerprint(payload: dict) -> str:
    return hashlib.sha256(canonical(payload)).hexdigest()

def build_expected(run_id: str, remote_request_id: str, calls: list) -> dict:
    spans = []
    for index, call in enumerate(calls):
        spans.append({
            "span_id": f"{run_id}-tool-{index:03d}",
            "run_id": run_id,
            "link": {"remote_request_id": remote_request_id},
            "tool_fingerprint": fingerprint({"name": call["name"], "args": call["args"]}),
            "result_fingerprint": fingerprint({"result": call.get("result")}),
        })
    return {
        "run_id": run_id,
        "remote_request_id": remote_request_id,
        "spans": spans,
    }

def score_allowed(expected: dict, ack: dict) -> bool:
    wanted = {span["span_id"] for span in expected["spans"]}
    wanted.add(expected["remote_request_id"])
    return wanted <= set(ack.get("accepted_ids", []))

def main() -> int:
    expected = json.loads(Path(sys.argv[1]).read_text())
    ack = json.loads(Path(sys.argv[2]).read_text())
    if not score_allowed(expected, ack):
        wanted = {span["span_id"] for span in expected["spans"]}
        wanted.add(expected["remote_request_id"])
        missing = sorted(wanted - set(ack.get("accepted_ids", [])))
        print(json.dumps({"status": "hold", "missing": missing}))
        return 2
    sequence = [span["tool_fingerprint"] for span in expected["spans"]]
    print(json.dumps({
        "status": "ready",
        "run_id": expected["run_id"],
        "sequence": sequence,
    }))
    return 0

if __name__ == "__main__":
    raise SystemExit(main())
Enter fullscreen mode Exit fullscreen mode

Build the expected file from the calls you actually made, before you export. The numbers in the fixture are placeholders, not measurements.

python3 - <<'PY'
import json
from ack_gate import build_expected

calls = [
    {"name": "read_file", "args": {"path": "src/app.py"}, "result": {"bytes": 1204}},
    {"name": "run_tests", "args": {"target": "tests/test_app.py"}, "result": {"failed": 0}},
]
doc = build_expected("run-1842", "req-9f3a", calls)
open("expected.json", "w").write(json.dumps(doc, indent=2) + "\n")
PY
python3 ack_gate.py expected.json ack.json
Enter fullscreen mode Exit fullscreen mode

Exit 2 means hold, and the JSON lists missing ids. Exit 0 means the fingerprint sequence may be compared with another ready run. A 202 from an async ingest URL is a queue ticket, not an ack file.

Persist the ack only from a body that names accepted ids. Include the remote request id in the same expected set as the tool spans. A tool-only ack still hides the model side, so the gate holds.

Span events are not a substitute for that body. An event can be sampled away while the span id still arrives, or the reverse. The score waits on ids the collector lists, not on events you hoped to attach.

Walk one failed export before you trust the script. Suppose the model span arrives and the second tool span does not, because the batch flushed early. The missing list contains run-1842-tool-001, and the sequence is withheld.

Retry the export with the same span ids. Replace ack.json with the new body, then run the gate again. Do not average the two acks, and do not mark the run partial-ready. Partial-ready is how a later diff blames the model for a span that never landed.

Idempotency belongs on the export, not on a fresh id per attempt. Reuse run-1842-tool-001 when you resend. A new id on each retry makes the expected set chase a moving target, and the ack can never catch it.

Where the free path fits

You can practice the split without standing up a private collector. MonkeyCode is the open-source project used here as the remote side: free model access for the model hop, and a free server option for the collector hop. Disclosure: This article was prepared as part of MonkeyCode's product outreach.

Both are availability claims, not a capacity plan. This article does not name models, token quotas, machine sizes, retention, or uptime. Those figures move, and a stale quota in a debug script is worse than no quota at all.

Read the current project documentation before you bind the loop to a limit. Do not paste an old token budget into the gate. If the docs and the runtime disagree, trust the runtime error and update the note.

Point the model hop at the free model access. Point the exporter at the free server, and send ids plus fingerprints rather than raw tool bodies. Keep expected.json beside the run directory so review does not depend on a dashboard session.

If the server schema cannot list accepted ids, do not fill them in locally. Extend the export until the ack is explicit, or leave the run on hold. A locally invented ack is just the manifest copied onto receipt paper.

The review step stays on the machine that ran the tools. When a ready sequence changes, open the local tool log for that run id and note the first hash that moved. The shared server is an index. It does not declare the agent fixed.

Who should not use this

The gate catches a missing export. It does not catch a wrong result that was exported in full. You still need an oracle on the fingerprint or on the local outcome.

It assumes the ack is specific and repeatable. If the server only says ok, the script holds forever, unless someone stubs accepted_ids with the expected set and voids the check. Do not stub it.

Do not use a free server option as audit storage, incident evidence, or a system of record. A convenience collector can drop or reset data. Do not send tool arguments that contain credentials, customer data, or private source.

Skip the gate if the model and the tools already share one process and one tracer. The extra files are ceremony you will ignore. Skip it if you cannot change the exporter to name accepted ids.

A gate without an honest ack is a new way to feel finished. Waiting costs whatever the batch interval costs. That wait is the method. A faster score on a partial trace is a confident miss.

Close on three files

Keep the expected set, the collector ack, and the fingerprint sequence. Diff sequences only across runs whose acks covered their sets. If the sequence is stable and the local oracle passed, record the run as unchanged.

If the sequence moved, inspect the local body for the first differing hash and stop. Do not keep walking the trace until you find a story that fits. The first hash delta is the debug pointer.

That is smaller than a claim that the agent is fixed. It is a claim you can recheck after the model label changes and the tool contract does not. A later swap should move the vendor string and leave the sequence alone, or the contract really changed.

If run notes already sit in the working tree, write the expected set there and export only ids and fingerprints. Recheck the current free model access and free server terms in the project docs. Let the ack file, not a status page, decide whether the run is scorable.

Top comments (0)