DEV Community

rene
rene

Posted on

You can't cut an LLM bill you don't measure: a 40-line stdlib per-call cost ledger in Python

Most teams find out what their LLM features cost when the monthly invoice arrives. By then you can't say which feature, prompt, or retry loop spent the money.

The fix is not a dashboard product. Start by writing one line per call to an append-only file, and tag each line with why the call happened. With that one file you can answer the 3 questions that matter:

  1. Which feature spends the most?
  2. Which prompt got more expensive after the last change?
  3. Is anything calling the model in a loop?

Here is a stdlib-only version you can drop into any Python project.

1. The ledger (append-only JSONL)

# cost_ledger.py - stdlib only
import json, time, os, functools

LEDGER = os.environ.get("LLM_LEDGER", "llm_costs.jsonl")

# USD per 1M tokens. Fill in from your provider's CURRENT pricing page -
# prices change, so never hard-code numbers you haven't checked.
PRICES = {
    # "model-name": (input_per_1m, output_per_1m),
}

def cost_usd(model, tokens_in, tokens_out):
    p_in, p_out = PRICES.get(model, (0.0, 0.0))
    return (tokens_in * p_in + tokens_out * p_out) / 1_000_000

def record(model, tokens_in, tokens_out, feature, ok=True, ms=0):
    row = {
        "ts": time.strftime("%Y-%m-%dT%H:%M:%S"),
        "model": model, "feature": feature,
        "in": tokens_in, "out": tokens_out,
        "usd": round(cost_usd(model, tokens_in, tokens_out), 6),
        "ok": ok, "ms": ms,
    }
    with open(LEDGER, "a", encoding="utf-8") as f:
        f.write(json.dumps(row) + "\n")
    return row
Enter fullscreen mode Exit fullscreen mode

Two design choices matter here:

  • Append-only JSONL. It survives crashes, you can tail -f it, and every tool can read it (jq, pandas, DuckDB, a spreadsheet).
  • A feature tag on every row. Without it you only know the total, and a total doesn't tell you what to fix.

2. Wrap your client call once

def tracked(feature):
    def deco(fn):
        @functools.wraps(fn)
        def wrapper(*args, **kwargs):
            t0 = time.time()
            try:
                resp = fn(*args, **kwargs)
            except Exception:
                record(kwargs.get("model", "?"), 0, 0, feature, ok=False,
                       ms=int((time.time() - t0) * 1000))
                raise
            usage = getattr(resp, "usage", None) or {}
            get = (lambda k: getattr(usage, k, None) or
                   (usage.get(k, 0) if isinstance(usage, dict) else 0))
            record(kwargs.get("model", "?"),
                   get("input_tokens") or get("prompt_tokens"),
                   get("output_tokens") or get("completion_tokens"),
                   feature, ms=int((time.time() - t0) * 1000))
            return resp
        return wrapper
    return deco
Enter fullscreen mode Exit fullscreen mode

Usage:

@tracked("ticket-summarizer")
def summarize(**kw):
    return client.messages.create(**kw)   # or chat.completions.create
Enter fullscreen mode Exit fullscreen mode

This handles both common usage shapes (input_tokens/output_tokens and prompt_tokens/completion_tokens). It also records failed calls, because a retry storm of failures still costs you latency and often tokens.

3. Answer the three questions

# report.py
import json, collections

rows = [json.loads(l) for l in open("llm_costs.jsonl", encoding="utf-8")]

by_feature = collections.Counter()
calls = collections.Counter()
for r in rows:
    by_feature[r["feature"]] += r["usd"]
    calls[(r["feature"], r["ts"][:16])] += 1   # calls per feature per minute

print("Spend by feature:")
for feat, usd in by_feature.most_common():
    print(f"  {feat:<25} ${usd:.4f}")

print("\nPossible loops (>30 calls in one minute):")
for (feat, minute), n in calls.items():
    if n > 30:
        print(f"  {feat} @ {minute}: {n} calls")
Enter fullscreen mode Exit fullscreen mode
  • Question 1 is the by_feature table.
  • Question 2: add a prompt_version field to record() and group by it. A prompt edit that doubles output tokens shows up the same day.
  • Question 3 is the per-minute counter. In an agent, a runaway loop almost always looks like dozens of calls to the same feature within one minute.

4. Turn it into a guard (optional)

Once the ledger exists, a budget check is a few lines:

def spent_today():
    today = time.strftime("%Y-%m-%d")
    with open(LEDGER, encoding="utf-8") as f:
        return sum(json.loads(l)["usd"] for l in f if today in l[:20])

if spent_today() > float(os.environ.get("LLM_DAILY_CAP", "5")):
    raise RuntimeError("Daily LLM budget reached - refusing call")
Enter fullscreen mode Exit fullscreen mode

Put that check at the top of wrapper and the ledger becomes a circuit breaker as well as a record.

What this doesn't do

  • It doesn't know prices. You fill in PRICES from your provider's pricing page and update it when prices change.
  • It's per-process. If several workers write to the same file, give each one its own ledger file, or move the rows into SQLite.
  • Cached or batch-discounted tokens need their own fields if your provider bills them differently.

If you'd rather not assemble and maintain this yourself, I packaged a fuller version (per-feature reports, CSV export, and budget alerts) as LLM Cost Tracker: https://renevibe76.gumroad.com/l/yknhue

The snippets above work fine on their own, though. Log every call and tag every call, and you can find what's driving the bill before the invoice arrives.

Top comments (0)