DEV Community

Sam Sun
Sam Sun

Posted on

Refuse the Next Attempt While the Trace File Is Shapeless

A second agent attempt is a cost, not a diagnosis, until the first trace file has a closed shape. The summary line can say the run finished, and that sentence still does not prove a root span exists, that each tool call has a child, or that an error carries a type you can group later. Treat the trace file as a packing list. If a label is missing, do not order another shipment.

Debug loops fail quietly when the artifacts do not share an id. The model returns a paragraph, a tool log shows a command, and the working tree may change, yet none of those three is a span. The next call then starts a new story, and you cannot tell whether it retried the same tool, skipped it, or invented a third path. A shape check sits in front of that call: it does not decide whether a patch is correct, only whether the run is comparable at all.

Shape is the packing list, not the cargo. A clerk can confirm a crate has a label, a weight, and a seal without opening every item. Presence of a trace id, unique span ids, parent links, a non-negative duration, a tool name, and an input hash is that seal. Whether the hash matches the file you meant to edit is a later question, and mixing the two is how a team spends a model call on a missing field.

Keep the contract small enough to fail for one reason: one file, one trace id, exactly one root whose parent is empty, and every other span pointing at a parent that also appears in the file. Span ids do not repeat, and end_ns is at least start_ns, which only checks that single span's own clock values rather than ordering across machines. Spans named with a tool. prefix carry tool.name and input.sha256, while a status of error carries an exception event or an error.type attribute, and the root carries run.id. A missing field is an exit code, not a warning scrolled past in CI.

Exit code 2 should stop the shell the way a failed test stops a deploy. A warning still lets the next step run, and the next step is the model call you were trying to protect. Exit code 0 means the file can be joined to a later attempt. It does not mean the agent fixed the bug, and it does not mean a remote collector has the same bytes.

The checker below is a proposed local tool: it has not been timed, and it is not a benchmark. It reads a simplified JSON export rather than OTLP protobuf. If your tracer already writes OTLP JSON, map spanId, parentSpanId, and startTimeUnixNano in a short adapter, then keep the gate itself dull. One failure reason per run is the point.

#!/usr/bin/env python3
"""Proposed trace-shape gate. Unexecuted example; adapt names to your exporter."""
import hashlib, json, sys

def fail(msg: str) -> None:
    print(f"shape: {msg}", file=sys.stderr)
    raise SystemExit(2)

def main() -> None:
    path = sys.argv[1] if len(sys.argv) > 1 else "run.json"
    with open(path, encoding="utf-8") as fh:
        doc = json.load(fh)
    spans = doc.get("spans")
    if not isinstance(spans, list) or not spans:
        fail("spans missing")
    ids = [s.get("span_id") for s in spans]
    if any(not i for i in ids) or len(ids) != len(set(ids)):
        fail("span_id missing or duplicated")
    roots = [s for s in spans if not s.get("parent_span_id")]
    if len(roots) != 1:
        fail(f"expected 1 root, found {len(roots)}")
    if not doc.get("trace_id"):
        fail("trace_id missing")
    known = set(ids)
    for s in spans:
        try:
            if int(s["end_ns"]) < int(s["start_ns"]):
                fail(f"{s['span_id']} negative duration")
        except (KeyError, TypeError, ValueError):
            fail(f"{s.get('span_id')} bad timestamps")
        parent = s.get("parent_span_id")
        if parent and parent not in known:
            fail(f"{s['span_id']} parent not in file")
        attrs = s.get("attributes") or {}
        if s is roots[0] and not attrs.get("run.id"):
            fail("root missing run.id")
        if str(s.get("name", "")).startswith("tool."):
            if not attrs.get("tool.name") or not attrs.get("input.sha256"):
                fail(f"{s['span_id']} tool span missing name or input hash")
        if s.get("status") == "error":
            events = [e.get("name") for e in (s.get("events") or [])]
            if "exception" not in events and not attrs.get("error.type"):
                fail(f"{s['span_id']} error has no type")
    digest = hashlib.sha256(",".join(ids).encode()).hexdigest()[:12]
    print(f"shape_ok trace={doc['trace_id']} spans={len(spans)} id_hash={digest}")

if __name__ == "__main__":
    main()
Enter fullscreen mode Exit fullscreen mode

Commit a fixture next to the script so the gate can be tested without an agent. The root needs run.id, and the tool span needs a hex input.sha256, status error, and error.type set to something like timeout. python3 trace_shape.py fixture.json should print one shape_ok line and exit 0; strip input.sha256 into a copy and expect exit 2. If the broken copy still exits 0, the gate is decorative and should not sit in front of any model call.

python3 trace_shape.py fixture.json; echo "pass:$?"
python3 - <<'PY'
import json
d = json.load(open("fixture.json", encoding="utf-8"))
d["spans"][1]["attributes"].pop("input.sha256", None)
json.dump(d, open("/tmp/broken.json", "w"), indent=2)
PY
python3 trace_shape.py /tmp/broken.json; echo "broken:$?"
Enter fullscreen mode Exit fullscreen mode

The fixture itself can stay under thirty lines. Two spans are enough to prove the root rule and the tool rule. Add a third span later only when you have a second tool. Extra spans that exist only to make the file look serious will train you to ignore the checker.

{
  "trace_id": "4f0c9a11c0de4b0a9c11aa00bb11cc22",
  "spans": [
    {
      "span_id": "aa",
      "parent_span_id": null,
      "name": "agent.run",
      "start_ns": 1000,
      "end_ns": 9000,
      "status": "ok",
      "attributes": {"run.id": "run-001"},
      "events": []
    },
    {
      "span_id": "bb",
      "parent_span_id": "aa",
      "name": "tool.read_file",
      "start_ns": 1500,
      "end_ns": 4000,
      "status": "error",
      "attributes": {
        "tool.name": "read_file",
        "input.sha256": "aaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaa",
        "error.type": "timeout"
      },
      "events": [{"name": "exception"}]
    }
  ]
}
Enter fullscreen mode Exit fullscreen mode

Store the raw tool argument outside the span, keyed by that hash, if you still need the body for a local replay. The gate only checks that the hash field is present and does not recompute it, because the body is not in this file. Recomputation belongs in a different script so a redaction change is not mistaken for an agent change. The id_hash printed on success is only a fingerprint of span ids in file order, useful in a CI log, and it will change if you sort spans differently.

A behavioral diff is a second step, and it should wait until two files pass, at which point you compare names, status codes, and input.sha256 values. A name that exists only in the second file is a new tool call, and a hash that changes under the same tool name is a new argument. Comparing two files that failed the gate produces a neat table of absences, which feels like analysis and is not. This gate stops before that comparison on purpose.

Skip the approach when a collector already enforces the same schema and the agent cannot emit a local file, because a weaker second gate will drift from the first. Skip it for sampled production traffic too: a sampler that drops tool children will make the checker fail runs you chose to thin out, and the exit code will blame instrumentation for a sampling decision. Skip it if exit code 0 would be shown to anyone as proof the working tree is correct, since a well-shaped trace can describe a wrong edit. Skip it if you cannot keep even a hash of tool arguments, because that hash is still a join key and a join key needs an owner and a retention rule.

Disclosure: This article was prepared as part of MonkeyCode's product outreach. Free model access matters here only as the budget a shapeless retry would waste, and the free server option matters only as a later export target after trace_shape.py has already passed on a local file, so collector setup and agent behavior are not debugged in the same hour. No model name, quota, hardware size, or retention period is stated, because none of those were verified for this draft; if that server is unreachable, keep the JSON, since a remote batch that never arrives neither creates nor erases a local pass. Once the file is green, that free access is a reasonable place for the next attempt, as long as the checked file stays the source of record.

The loop stays short on purpose: emit one JSON file, run the shape script, and repair instrumentation until the broken fixture exits 2 and the real file exits 0. Then, and only then, spend another model call. The summary line can wait. The packing list cannot.

Top comments (0)