A reviewer opens a failed agent run at 02:14.
The status line says the model call was rejected.
Tool six never returned a file diff at all.
The log stores one error string and nothing else.
It lacks a split of input tokens and output tokens.
Nobody can name the span that spent the budget.
The hole that wastes the next hour
Failed agent runs often keep the final error string only.
They omit the running token total at each step.
A retry loop then looks like one red step.
A large file read can look like a planner defect.
A planner defect can look like a quota failure.
Those faults need different fixes and different owners.
This note adds one record type to the trace.
Each model span and each tool span writes one row.
You join those rows by parent span id.
The sample below is an unexecuted local example.
Run it on your machine before you trust the totals.
The fixture numbers are not a hosted model benchmark.
What the ledger must hold
Store one ledger row for every span in the run.
Keep the field set small so writers do not skip rows.
Add a run header before the first span row.
Header fields:
-
run_id: stable id for this attempt -
unit:input_plus_outputor the vendor unit name -
local_soft_limit: your ceiling, not a vendor promise -
local_hard_limit: stop line for this process -
prompt_storage:hash_onlyin this design
Row fields:
-
span_id: unique inside the run -
parent_id: caller id, empty at the root -
seq: integer order inside the run -
kind:modelortool -
name: model call label or tool name -
input_tokens: integer, or null if unknown -
output_tokens: integer, or null if unknown -
running_total: sum of known tokens so far -
status:ok,error,budget_warn, orbudget_stop -
ts_mono_ms: monotonic milliseconds inside one process
Do not store raw prompts in this ledger file.
Prompts can hold secrets and entire file bodies.
Keep a prompt hash only if you already hash prompts.
Unknown usage must stay null, not zero.
Zero means the host reported a measured zero.
Null means the host did not report usage at all.
A recorder you can execute locally
This script writes JSONL and never opens a socket.
It is a proposal you should run before adoption.
Replace the fixture limits with your real local ceiling.
#!/usr/bin/env python3
"""Unexecuted example: per-span token ledger. Not a benchmark."""
import json
from pathlib import Path
class TokenLedger:
def __init__(self, path, run_id, soft_limit, hard_limit):
self.path = Path(path)
self.run_id = run_id
self.soft_limit = soft_limit
self.hard_limit = hard_limit
self.total = 0
self.seq = 0
self.rows = []
header = {
"run_id": run_id,
"unit": "input_plus_output",
"local_soft_limit": soft_limit,
"local_hard_limit": hard_limit,
"prompt_storage": "hash_only",
}
self.path.write_text(
json.dumps({"header": header}) + "\n",
encoding="utf-8",
)
def add(self, span_id, parent_id, kind, name,
input_tokens, output_tokens, status):
self.seq += 1
known = 0
for value in (input_tokens, output_tokens):
if value is not None:
known += int(value)
self.total += known
if self.total >= self.hard_limit and status == "ok":
status = "budget_stop"
elif self.total >= self.soft_limit and status == "ok":
status = "budget_warn"
row = {
"run_id": self.run_id,
"span_id": span_id,
"parent_id": parent_id,
"seq": self.seq,
"kind": kind,
"name": name,
"input_tokens": input_tokens,
"output_tokens": output_tokens,
"running_total": self.total,
"status": status,
}
self.rows.append(row)
with self.path.open("a", encoding="utf-8") as handle:
handle.write(json.dumps(row) + "\n")
return row
def by_parent(self):
grouped = {}
for row in self.rows:
key = row["parent_id"] or "ROOT"
spent = 0
for value in (row["input_tokens"], row["output_tokens"]):
if value is not None:
spent += int(value)
grouped[key] = grouped.get(key, 0) + spent
ranked = sorted(
grouped.items(),
key=lambda item: item[1],
reverse=True,
)
return ranked
if __name__ == "__main__":
ledger = TokenLedger("usage.jsonl", "run-demo", 800, 1000)
ledger.add("s1", "", "model", "plan", 120, 40, "ok")
ledger.add("s2", "s1", "tool", "read_file", 0, 0, "ok")
ledger.add("s3", "s1", "model", "patch", 700, 180, "ok")
print(json.dumps(ledger.by_parent()))
The demo parent s1 collects the child spend.
The patch call crosses the hard fixture limit.
Its status becomes budget_stop inside the local writer.
Those token counts are fixtures for the assert.
They do not describe any live vendor model.
Checks that must pass before blame
File assert
Run this consistency check after every recorded attempt.
It refuses a file whose running total drifts.
python3 - << 'PY'
import json
total = 0
prev = 0
seen = set()
for line in open("usage.jsonl", encoding="utf-8"):
obj = json.loads(line)
if "header" in obj:
assert obj["header"]["unit"] == "input_plus_output"
continue
spent = 0
for key in ("input_tokens", "output_tokens"):
if obj[key] is not None:
spent += obj[key]
total += spent
assert obj["running_total"] == total, obj["span_id"]
assert obj["seq"] == prev + 1
assert obj["span_id"] not in seen
seen.add(obj["span_id"])
prev = obj["seq"]
print("rows", prev, "known_tokens", total)
PY
If an assert fires, stop and fix the writer.
Do not open the model prompt until the file is consistent.
A drifted total will point you at the wrong span.
Join rules
Four join rules keep the file boring and useful.
Apply them before anyone debates model quality.
- Each
span_idappears once inside the run. - Each non-root
parent_idmatches an earlier span. -
running_totalequals prior known spend plus this row. - A
budget_stoprow is the last row in the file.
A later tool row after budget_stop is a writer bug.
The agent should have halted before that tool started.
Record the bug, then fix the stop path first.
Debug loop for a budget death
Five steps
Use the same five steps whenever a run dies on tokens.
Do not skip ahead to a prompt edit.
- Read the last row and note status, name, and total.
- Group known spend by
parent_idand rank the parents. - Open the top parent and list each child name.
- Mark the child with the largest known sum as suspect.
- Replay only that child with a smaller input body.
One change per replay
Change only one variable on each replay pass.
Keep the tool set fixed while you shrink the input.
Then keep the input fixed while you cap retries.
If the new total stays under your local hard limit, stop.
You found a spend path, not a model quality verdict.
Quality checks belong in a separate harness and note.
How to read the stop table
| Signal | Likely cause | Next action |
|---|---|---|
Last row is budget_stop on a model span |
That call or its parent loop hit the local cap | Shrink the input or cap the loop |
| Many model rows share one parent id | Retry storm or planner loop | Set a retry cap before the next run |
| Tool rows show zero, then one huge model row | Tool output was copied into the next prompt | Truncate tool output before the model call |
Total jumps while seq has gaps |
The writer dropped rows | Repair the writer, then rerun the same task |
| Vendor error says quota but local total is low | Units or hidden host tokens disagree | Reconcile units before you blame the prompt |
Pick one row in that table per incident.
Do not apply every action in a single replay.
Mixed changes hide the span that actually spent the budget.
Shared traces and free model access
Disclosure: This article was prepared as part of MonkeyCode's product outreach.
The operator describes MonkeyCode as an open-source project.
The same briefing cites free model access and a free server option.
This note states no token quota, model list, hardware size, or expiry.
Those details change, and no primary source is attached here.
Check the project page on the day you plan a run.
A free server helps only as shared storage for redacted traces.
Upload the usage file and the run header only.
Leave raw prompts and raw tool bodies on the machine that created them.
Teammates need the span name and the running total.
They do not need secrets that rode along in a file read.
Strip those fields before the file leaves the laptop.
The debug loop itself runs without a hosted model.
Prove the join asserts on fixture rows first.
Point a live call at the ledger only after those asserts pass.
Limits that stay in force
This ledger cannot see billing the host refuses to return.
A null usage field is a gap, not a measured zero.
Say that gap in the incident note so others do not invent a number.
Parent joins fail when the runtime omits parent ids.
Fall back to seq order and stop drawing a tree.
A flat list is weaker, and it is still honest.
Monotonic time is valid inside one process only.
Do not sort rows from two hosts by wall clock.
Clock skew will invent overlaps that never happened.
The soft limit and hard limit in the sample are fixtures.
Replace both with the ceiling you chose for this run.
If the vendor has not published a cap, label your cap as local.
Treat this file as a debug aid only.
It is not an invoice and not a legal usage record.
Do not cite it as billed spend in a customer report.
Who should skip the extra file
Skip the ledger if your runtime already exports billed usage per span.
A second counter will drift and start avoidable arguments.
Use the vendor export and document its unit in the run header.
Skip the upload if you cannot redact tool output first.
A shared server is the wrong store for credentials and private source.
Keep those runs on disk with restricted permissions.
Skip this design if you need finance-grade metering.
Local JSONL can drop rows when a process dies hard.
That gap is acceptable for debugging and unacceptable for billing.
One next run
Add the four join asserts to the last failed attempt.
Write the ledger beside the trace you already store.
If your team has a free shared server, upload the redacted ledger only.
Confirm the live access limits on the project page that same day.
Then replay one suspect span, and leave the rest of the agent untouched.
Top comments (0)