DEV Community

Riley Wang
Riley Wang

Posted on

Name the Span That Spent the Agent Budget

A reviewer opens a failed agent run at 02:14.
The status line says the model call was rejected.
Tool six never returned a file diff at all.

The log stores one error string and nothing else.
It lacks a split of input tokens and output tokens.
Nobody can name the span that spent the budget.

The hole that wastes the next hour

Failed agent runs often keep the final error string only.
They omit the running token total at each step.
A retry loop then looks like one red step.

A large file read can look like a planner defect.
A planner defect can look like a quota failure.
Those faults need different fixes and different owners.

This note adds one record type to the trace.
Each model span and each tool span writes one row.
You join those rows by parent span id.

The sample below is an unexecuted local example.
Run it on your machine before you trust the totals.
The fixture numbers are not a hosted model benchmark.

What the ledger must hold

Store one ledger row for every span in the run.
Keep the field set small so writers do not skip rows.
Add a run header before the first span row.

Header fields:

  • run_id: stable id for this attempt
  • unit: input_plus_output or the vendor unit name
  • local_soft_limit: your ceiling, not a vendor promise
  • local_hard_limit: stop line for this process
  • prompt_storage: hash_only in this design

Row fields:

  • span_id: unique inside the run
  • parent_id: caller id, empty at the root
  • seq: integer order inside the run
  • kind: model or tool
  • name: model call label or tool name
  • input_tokens: integer, or null if unknown
  • output_tokens: integer, or null if unknown
  • running_total: sum of known tokens so far
  • status: ok, error, budget_warn, or budget_stop
  • ts_mono_ms: monotonic milliseconds inside one process

Do not store raw prompts in this ledger file.
Prompts can hold secrets and entire file bodies.
Keep a prompt hash only if you already hash prompts.

Unknown usage must stay null, not zero.
Zero means the host reported a measured zero.
Null means the host did not report usage at all.

A recorder you can execute locally

This script writes JSONL and never opens a socket.
It is a proposal you should run before adoption.
Replace the fixture limits with your real local ceiling.

#!/usr/bin/env python3
"""Unexecuted example: per-span token ledger. Not a benchmark."""

import json
from pathlib import Path

class TokenLedger:
    def __init__(self, path, run_id, soft_limit, hard_limit):
        self.path = Path(path)
        self.run_id = run_id
        self.soft_limit = soft_limit
        self.hard_limit = hard_limit
        self.total = 0
        self.seq = 0
        self.rows = []
        header = {
            "run_id": run_id,
            "unit": "input_plus_output",
            "local_soft_limit": soft_limit,
            "local_hard_limit": hard_limit,
            "prompt_storage": "hash_only",
        }
        self.path.write_text(
            json.dumps({"header": header}) + "\n",
            encoding="utf-8",
        )

    def add(self, span_id, parent_id, kind, name,
            input_tokens, output_tokens, status):
        self.seq += 1
        known = 0
        for value in (input_tokens, output_tokens):
            if value is not None:
                known += int(value)
        self.total += known
        if self.total >= self.hard_limit and status == "ok":
            status = "budget_stop"
        elif self.total >= self.soft_limit and status == "ok":
            status = "budget_warn"
        row = {
            "run_id": self.run_id,
            "span_id": span_id,
            "parent_id": parent_id,
            "seq": self.seq,
            "kind": kind,
            "name": name,
            "input_tokens": input_tokens,
            "output_tokens": output_tokens,
            "running_total": self.total,
            "status": status,
        }
        self.rows.append(row)
        with self.path.open("a", encoding="utf-8") as handle:
            handle.write(json.dumps(row) + "\n")
        return row

    def by_parent(self):
        grouped = {}
        for row in self.rows:
            key = row["parent_id"] or "ROOT"
            spent = 0
            for value in (row["input_tokens"], row["output_tokens"]):
                if value is not None:
                    spent += int(value)
            grouped[key] = grouped.get(key, 0) + spent
        ranked = sorted(
            grouped.items(),
            key=lambda item: item[1],
            reverse=True,
        )
        return ranked

if __name__ == "__main__":
    ledger = TokenLedger("usage.jsonl", "run-demo", 800, 1000)
    ledger.add("s1", "", "model", "plan", 120, 40, "ok")
    ledger.add("s2", "s1", "tool", "read_file", 0, 0, "ok")
    ledger.add("s3", "s1", "model", "patch", 700, 180, "ok")
    print(json.dumps(ledger.by_parent()))
Enter fullscreen mode Exit fullscreen mode

The demo parent s1 collects the child spend.
The patch call crosses the hard fixture limit.
Its status becomes budget_stop inside the local writer.

Those token counts are fixtures for the assert.
They do not describe any live vendor model.

Checks that must pass before blame

File assert

Run this consistency check after every recorded attempt.
It refuses a file whose running total drifts.

python3 - << 'PY'
import json
total = 0
prev = 0
seen = set()
for line in open("usage.jsonl", encoding="utf-8"):
    obj = json.loads(line)
    if "header" in obj:
        assert obj["header"]["unit"] == "input_plus_output"
        continue
    spent = 0
    for key in ("input_tokens", "output_tokens"):
        if obj[key] is not None:
            spent += obj[key]
    total += spent
    assert obj["running_total"] == total, obj["span_id"]
    assert obj["seq"] == prev + 1
    assert obj["span_id"] not in seen
    seen.add(obj["span_id"])
    prev = obj["seq"]
print("rows", prev, "known_tokens", total)
PY
Enter fullscreen mode Exit fullscreen mode

If an assert fires, stop and fix the writer.
Do not open the model prompt until the file is consistent.
A drifted total will point you at the wrong span.

Join rules

Four join rules keep the file boring and useful.
Apply them before anyone debates model quality.

  1. Each span_id appears once inside the run.
  2. Each non-root parent_id matches an earlier span.
  3. running_total equals prior known spend plus this row.
  4. A budget_stop row is the last row in the file.

A later tool row after budget_stop is a writer bug.
The agent should have halted before that tool started.
Record the bug, then fix the stop path first.

Debug loop for a budget death

Five steps

Use the same five steps whenever a run dies on tokens.
Do not skip ahead to a prompt edit.

  1. Read the last row and note status, name, and total.
  2. Group known spend by parent_id and rank the parents.
  3. Open the top parent and list each child name.
  4. Mark the child with the largest known sum as suspect.
  5. Replay only that child with a smaller input body.

One change per replay

Change only one variable on each replay pass.
Keep the tool set fixed while you shrink the input.
Then keep the input fixed while you cap retries.

If the new total stays under your local hard limit, stop.
You found a spend path, not a model quality verdict.
Quality checks belong in a separate harness and note.

How to read the stop table

Signal Likely cause Next action
Last row is budget_stop on a model span That call or its parent loop hit the local cap Shrink the input or cap the loop
Many model rows share one parent id Retry storm or planner loop Set a retry cap before the next run
Tool rows show zero, then one huge model row Tool output was copied into the next prompt Truncate tool output before the model call
Total jumps while seq has gaps The writer dropped rows Repair the writer, then rerun the same task
Vendor error says quota but local total is low Units or hidden host tokens disagree Reconcile units before you blame the prompt

Pick one row in that table per incident.
Do not apply every action in a single replay.
Mixed changes hide the span that actually spent the budget.

Shared traces and free model access

Disclosure: This article was prepared as part of MonkeyCode's product outreach.
The operator describes MonkeyCode as an open-source project.
The same briefing cites free model access and a free server option.

This note states no token quota, model list, hardware size, or expiry.
Those details change, and no primary source is attached here.
Check the project page on the day you plan a run.

A free server helps only as shared storage for redacted traces.
Upload the usage file and the run header only.
Leave raw prompts and raw tool bodies on the machine that created them.

Teammates need the span name and the running total.
They do not need secrets that rode along in a file read.
Strip those fields before the file leaves the laptop.

The debug loop itself runs without a hosted model.
Prove the join asserts on fixture rows first.
Point a live call at the ledger only after those asserts pass.

Limits that stay in force

This ledger cannot see billing the host refuses to return.
A null usage field is a gap, not a measured zero.
Say that gap in the incident note so others do not invent a number.

Parent joins fail when the runtime omits parent ids.
Fall back to seq order and stop drawing a tree.
A flat list is weaker, and it is still honest.

Monotonic time is valid inside one process only.
Do not sort rows from two hosts by wall clock.
Clock skew will invent overlaps that never happened.

The soft limit and hard limit in the sample are fixtures.
Replace both with the ceiling you chose for this run.
If the vendor has not published a cap, label your cap as local.

Treat this file as a debug aid only.
It is not an invoice and not a legal usage record.
Do not cite it as billed spend in a customer report.

Who should skip the extra file

Skip the ledger if your runtime already exports billed usage per span.
A second counter will drift and start avoidable arguments.
Use the vendor export and document its unit in the run header.

Skip the upload if you cannot redact tool output first.
A shared server is the wrong store for credentials and private source.
Keep those runs on disk with restricted permissions.

Skip this design if you need finance-grade metering.
Local JSONL can drop rows when a process dies hard.
That gap is acceptable for debugging and unacceptable for billing.

One next run

Add the four join asserts to the last failed attempt.
Write the ledger beside the trace you already store.
If your team has a free shared server, upload the redacted ledger only.

Confirm the live access limits on the project page that same day.
Then replay one suspect span, and leave the rest of the agent untouched.

Top comments (0)