DEV Community

rene
rene

Posted on

My AI agent script almost burned through a $25 budget in one afternoon. Here's the 10-line Python fix.

I run a small swarm of autonomous AI agents — they write code, post content, trade on paper markets, and ship small products. The whole operation runs on a budget: $25 of LLM API credits, total, for the month.

Last week one agent hit a bug: a retry loop with no upper bound. If it had kept going, it would have chewed through that entire $25 cap in under an hour — one bad response, retried forever, each retry costing a few cents.

It didn't, because every call in this swarm goes through a 10-line guard first. Here it is, stdlib only, no dependencies:

import time, json, os

BUDGET_USD = 25.0
STATE_FILE = "budget_state.json"

def _load():
    if os.path.exists(STATE_FILE):
        with open(STATE_FILE) as f:
            return json.load(f)
    return {"spent": 0.0, "calls": 0}

def _save(state):
    with open(STATE_FILE, "w") as f:
        json.dump(state, f)

def guarded_call(cost_estimate, fn, *args, **kwargs):
    state = _load()
    if state["spent"] + cost_estimate > BUDGET_USD:
        raise RuntimeError(
            f"Budget guard: ${state['spent']:.2f} already spent, "
            f"next call (${cost_estimate:.3f}) would exceed the ${BUDGET_USD} cap. Stopping."
        )
    result = fn(*args, **kwargs)
    state["spent"] += cost_estimate
    state["calls"] += 1
    _save(state)
    return result
Enter fullscreen mode Exit fullscreen mode

That's it. Instead of calling your LLM/API client directly, you wrap the call:

response = guarded_call(0.012, call_llm, prompt="summarize this")
Enter fullscreen mode Exit fullscreen mode

If the running total would cross the cap, it raises before the network call happens — not after you get the bill. No silent runaway loop, no 3am Slack alert about a $400 invoice.

Two things I'd add if you're shipping this for real:

  1. A loop breaker — count consecutive calls in the same run and bail after N, even if budget remains (a stuck loop can spam fast and still blow past your intended pace).
  2. A daily reset — reset spent on a rolling 24h window instead of letting it accumulate forever, so one bad day doesn't block tomorrow's legitimate work.

We packaged the fuller version (budget cap + loop breaker + run log) as a free, stdlib-only download if you want the ready-made version instead of writing your own: https://renevibe76.gumroad.com/l/kwgtni

Has a runaway agent or script ever surprised you with an API bill — or did a guard like this catch it in time for you? What's your stop condition?

Top comments (1)

Collapse
 
chrissellers profile image
Chris Sellers •

Two holes I'd look at in this version, both in the snippet as posted.

First, spend is only recorded after fn returns. If the call raises (a timeout, a 5xx, a parse error on a half read response), nothing is added, but the provider may still bill for it. A loop of failing retries is exactly the runaway case, and this guard lets it run forever. Reserving the estimate before the call, and refunding it only when you know the request never left, closes that.

Second, with a swarm, two agents can _load() the same total, both pass the check, and both write back, so one charge disappears. A file lock around load, check and save, or a small SQLite table with an atomic UPDATE ... WHERE spent + ? <= cap, fixes it.

On your stop condition question: I'd want a cap on consecutive failures per run on top of the dollar cap, since a broken call tends to fail fast and cheap right up until it doesn't.