A coding agent writes a three-file patch during a long run. The remote inference log records HTTP 200 for that turn. The local tool span records exit code 1 for the patch.
The agent summary still reports that the edit landed cleanly. No shared identifier connects those three records together. The failure sits in the gap between the two logs.
Recent community posts spend their energy on model choice. This note does not rank models or repeat those arguments. It asks whether both logs can name the same attempt.
Two clocks, two id spaces
Local traces usually name tools, arguments, and exit codes. Remote logs usually name model calls, status, and latency. Each side keeps its own clock and its own identifier space.
A retry writes a second model call under a new request id. The tool span can still point at the first attempt only. A file diff will not show that the join itself broke.
This split appears on long agent runs with several tools. It appears again when the runner and the server sit far apart. Delay makes the timelines look unrelated even when they match.
What you can reuse
This note gives you a join record and a small checker. It also gives you four mismatch classes and a debug loop. The code is a proposed example and was not executed here.
It is not a benchmark of any model or any server. It does not quote a token allowance or a hardware spec. Treat fixture counts as labels for the method, not as measurements.
Fields that must travel together
Store one join row for every model turn and every tool span.
-
trace_ididentifies the whole agent run from start to finish. -
turn_ididentifies one model request and the matching response. -
tool_call_ididentifies one tool call, or stays null on model rows. -
parent_turn_idnames the model turn that requested that tool. -
attemptis an integer that starts at 1 for each logical turn. -
sideis eithermodelortool, and no third value is valid. -
server_request_idstores the remote id, or null when that id is absent. -
statusis a short code such asokorerror. -
local_mono_nsstores local monotonic time, counted in nanoseconds.
Do not store raw prompts or tool output in this join file. Store a content hash when you need an equality check later. Keep secrets out of the file before you export it anywhere.
Validate before you join
A checker that guesses parents will hide emitter bugs. Fail closed when a required field is missing or contradictory. Do not repair a row by matching the nearest timestamp.
# Proposed example. Not executed in this draft.
from dataclasses import dataclass
@dataclass(frozen=True)
class JoinRow:
trace_id: str
turn_id: str
tool_call_id: str | None
parent_turn_id: str | None
attempt: int
side: str
server_request_id: str | None
status: str
local_mono_ns: int
def validate(row: JoinRow) -> list[str]:
errors: list[str] = []
if not row.trace_id or not row.turn_id:
errors.append("missing_ids")
if row.attempt < 1:
errors.append("bad_attempt")
if row.side not in {"model", "tool"}:
errors.append("bad_side")
if row.side == "tool" and not row.parent_turn_id:
errors.append("tool_without_parent_turn")
if row.side == "tool" and not row.tool_call_id:
errors.append("tool_without_tool_id")
if row.side == "model" and row.tool_call_id is not None:
errors.append("model_row_has_tool_id")
if row.local_mono_ns < 0:
errors.append("bad_clock")
return errors
Run validate on every row before the join step starts. Stop the report when any row returns a schema error. A partial join over dirty input looks precise and is not.
Classify the gap
Index model rows by trace id, turn id, and attempt. Look up each tool span by its parent turn id. Return one class string, and do not attach a score.
# Proposed example. Not executed in this draft.
def index_models(rows: list[JoinRow]) -> dict[tuple[str, str, int], JoinRow]:
indexed = {}
for row in rows:
if row.side != "model":
continue
indexed[(row.trace_id, row.turn_id, row.attempt)] = row
return indexed
def classify(tool: JoinRow, models: dict) -> str:
if tool.side != "tool":
return "not_a_tool_row"
if not tool.parent_turn_id:
return "orphan_tool"
key = (tool.trace_id, tool.parent_turn_id, tool.attempt)
parent = models.get(key)
if parent is None:
return "missing_model_row"
if parent.server_request_id and tool.server_request_id:
if parent.server_request_id != tool.server_request_id:
return "request_id_mismatch"
if parent.status != "ok" and tool.status == "ok":
return "tool_ok_after_model_error"
return "joined"
Empty parent ids should die in validation, not in classify. The orphan class remains for checkers that skipped that gate. The checker does not rank models and does not score patches.
It only reports whether the two logs agree on identity. A joined pair can still contain a wrong edit. Agreement is a precondition, not a verdict on the patch.
Read the four classes
| Class | Local evidence | Remote evidence | First check |
|---|---|---|---|
orphan_tool |
tool span, empty parent | nothing required | validation should already have failed |
missing_model_row |
parent id is set | no matching model row | dropped log, wrong id, or sampling |
request_id_mismatch |
server id A on the tool | server id B on the model | a retry wrote a new call |
tool_ok_after_model_error |
tool status is ok
|
model status is not ok
|
summary trusted the wrong side |
Read the table from the class column toward the check column. Fix the emitter before you rewrite the prompt or the tool. A missing remote row is not proof that the call never ran.
Fixture, not a measurement
Use a tiny fixture so the four classes stay reviewable. Label it as synthetic data inside the test file itself. Do not paste these counts into a status report as results.
- Build two model rows with attempts 1 and 2 under one trace.
- Build three valid tool rows: one mismatch, one missing parent, one join.
- Expect
request_id_mismatch,missing_model_row, andjoined. - Expect zero schema errors, and fail the test if any appear.
# Proposed fixture. Not executed in this draft.
def test_fixture_classes() -> None:
rows = [
JoinRow("run_1844", "turn_a", None, None, 1, "model", "req_1", "ok", 10),
JoinRow("run_1844", "turn_a", None, None, 2, "model", "req_2", "ok", 20),
JoinRow("run_1844", "tool_row_1", "tool_1", "turn_missing", 1, "tool", "req_x", "error", 30),
JoinRow("run_1844", "tool_row_2", "tool_2", "turn_a", 1, "tool", "req_9", "ok", 12),
JoinRow("run_1844", "tool_row_3", "tool_3", "turn_a", 2, "tool", "req_2", "ok", 21),
]
assert all(validate(row) == [] for row in rows)
models = index_models(rows)
got = [classify(row, models) for row in rows if row.side == "tool"]
assert sorted(got) == ["joined", "missing_model_row", "request_id_mismatch"]
The fixture expects those mismatch classes to appear. The gate command below is for a later run, after the emitter fix. Do not point the gate at this fixture and call the failure a product bug.
# Proposed commands. Paths are examples only.
python join_check.py --local runs/trace.jsonl --remote runs/inference.jsonl --out runs/join-report.json
python join_check.py --local runs/trace.jsonl --remote runs/inference.jsonl --expect-zero orphan_tool,request_id_mismatch
The second command is a regression gate for the emitter. It should fail while those two mismatch classes still appear. Keep missing model rows out of that gate until retention is known.
Report shape
Write the checker output as a small JSON object. Mark synthetic runs so nobody quotes them as production data. Include the input file hash so a later replay can detect edits.
{
"trace_id": "run_1844",
"fixture": true,
"schema_errors": 0,
"input_sha256": "replace-with-real-hash",
"counts": {
"joined": 1,
"orphan_tool": 0,
"missing_model_row": 1,
"request_id_mismatch": 1,
"tool_ok_after_model_error": 0
}
}
Six steps you can rerun
- Export local tool spans to one JSONL file you can hash.
- Export remote inference rows that carry the same turn id.
- Validate every row, and stop on the first schema error.
- Classify each tool span, then write class counts to a new file.
- Open only the mismatch classes, and leave joined pairs for later.
- Patch the emitter, replay the same inputs, and confirm the drop.
Keep both input files immutable during that replay. Write each report to a new path with the trace hash in the name. That habit stops a later edit from rewriting the evidence.
Free server, same join contract
Disclosure: This article was prepared as part of MonkeyCode's product outreach.
The operator describes MonkeyCode as an open-source agent project. The same brief says free model access and a free server option exist. This draft does not state a token cap or a model list.
It also omits machine size and any uptime promise. Those access details change, and stale numbers mislead planning. Confirm the current limits in the project documentation before a long run.
Do not treat a blog post, including this one, as the pricing source. A free server helps this loop in one narrow way. You can place the agent loop on that server instead of the laptop.
You still keep the local join file under your control. That file remains the record you hold when remote retention is short. Free access does not remove the two-log problem.
A remote log can rotate before you export it. A rate limit can create attempt 2 with no matching tool span. Record attempt on both sides or you will join the wrong call.
Echo the server request id into the local row when one is returned. If the server cannot echo the turn id, stop and fix that contract. Use the free option as a second environment, not as an oracle.
Run the same schema in both places and diff the class counts. Diff those class-count reports, not the product pages. A lower count means the emitter improved, not that the model improved.
Limits you should keep visible
- If neither side logged the call, this checker cannot recover it.
- This checker does not prove the patch contents or the exit result.
- Use clocks as diagnostic notes, and never join rows by time.
- A short prompt hash can collide across different inputs.
- Hash the full canonical input, or skip the hash field entirely.
- A null server request id means the value was not captured.
- That null does not mean the call stayed on the local machine.
- Sampled remote logs create false missing-model-row hits.
- Mark sampling in the report header before anyone triages them.
Do not upload raw prompts just to fill an empty join field. Redact secrets before any export to a shared or free server. If redaction is unclear, keep the join file on the runner.
Who should skip this method
Skip it when one process already writes one complete log. Skip it when you cannot change the emitter that mints ids. Skip it when the question is model quality rather than trace integrity.
A fixed task harness answers quality questions better than this checker. This checker answers whether the two logs describe the same attempt. Mixing those questions produces a confident report about the wrong fault.
Also skip the remote half when the server log is heavily sampled. Say that in the report so a reader does not chase empty rows. Local validation can still run alone in that sampled case.
One pass on a failed trace
Pick one failed run that already has a local JSONL trace. Add the turn id and the parent turn id at the emitter. Classify once, and keep that report beside the trace as a baseline.
If you use MonkeyCode's described free server option, copy the request id. Put that id into the same local join row before you export. Check current access limits in the project docs before the second run.
Repeat the six steps in that second environment. Stop when the gated classes hit zero, then reread the patch. The join is ready only when the same attempt is named on both sides.
Top comments (0)