DEV Community

Ken @ Flaming Ape
Ken @ Flaming Ape

Posted on

Why a monthly spend cap won't stop a runaway AI agent loop

If you run AI agents, you've probably set a monthly spend limit and felt a little better about it. Fair enough. But a monthly cap answers "how much can I lose this month?" It doesn't answer "how do I stop this agent right now?" Those are two different problems, and runaway loops live in the gap between them.

Here's what I've learned about that gap, plus a checklist for any stack.

Monthly and invoice caps fire too late

A monthly cap is a ceiling on the whole account. When it trips, the money's already gone. It usually stops everything at once, the stuck agent and the healthy ones alike.

Billing alerts have the same problem. Usage data is often delayed, and an email that lands after you've gone to bed stops nothing. A cap high enough to never block normal work is high enough to let one bad loop chew through plenty.

The better question is narrower. How much should this one agent, on this one task, be allowed to spend before something pauses it?

Loops and retries multiply cost quietly

Agent loops rarely look dramatic. They look like normal work that never finishes. A test fails, the agent "fixes" it, the test fails the same way, and around it goes again. Or a tool call returns an error and the agent retries with almost the same input. Meanwhile the context keeps growing, so every call costs more than the last.

None of that throws an exception. As far as the API knows, every call succeeded. The cost climbs because nothing is counting turns, repeated actions or spend per task. Repeat detection deserves its own control, since it catches "same thing, again" even when each attempt is cheap.

Per-agent budgets, and a stop that pauses one agent, not all

Give each agent (or each task, or each API key) its own budget, sized for the job. A code-review bot and a long-running refactor agent shouldn't share one pool.

When a budget or a loop limit trips, the right move is to pause that one agent and leave the rest alone. Paused beats killed. You can see what it was doing, fix the prompt or raise the limit, and resume without losing work.

That means tagging every request with its agent and task. If you can't tell who spent it, you can't limit it.

Rate limits: queue and back off, don't just fail

Rate limits are a different failure from overspending, but they make loops worse. An agent that gets a 429 and retries right away can turn one rate limit into a retry storm.

I'd rather handle rate limits underneath the agent. Queue the request, honor any Retry-After header, back off with jitter, and hand the agent a response when it's ready. If the wait runs too long, return a clear, specific error, not a vague one that invites another retry.

Keep the meter outside the agent

If your budget lives in the system prompt ("you have $5, stop when you're done"), that's a suggestion, not a limit. Models miscount, lose track in long contexts, and get steered by things they read, like a README or a tool's output.

Enforcement belongs in something the agent can't edit or argue with: a proxy in front of the model API, your own wrapper around the client, or key limits at the provider. Tell the agent its budget. Just don't make it the enforcer.

A checklist you can use today

Most of this needs no new tools:

  • Separate API keys per agent or project, with provider-side limits where available.
  • An iteration cap on every agent loop. Most frameworks have one (max_iterations, max turns). Set it on purpose.
  • Per-task spend tracking, with tokens and cost logged per request and tagged by agent and task.
  • Repeat detection. Flag the same tool call or near-identical prompt several times in a row.
  • Rate-limit handling in one place, with backoff and Retry-After support.
  • Pause, don't kill. Stopping an agent should keep its state.
  • Enforcement outside the prompt that the agent can't override.
  • Keep the monthly cap as the last line of defense, not the first.

Where I'm at

Full disclosure: I run Flaming Ape LLC, a small veteran-owned company in Texas, and I'm building Gorilla Warden, a free, open-source, self-hosted core that sits in front of AI agents to handle per-agent budgets and rate-limit queuing. It isn't released yet. The core safety architecture is implemented and passing focused tests. Integration work, adversarial testing and an independent review are still ahead, and I'm aiming for a free public beta on GitHub in November 2026. I'm building it with AI help too (Claude Code, Codex and Grok), which is part of why this problem is on my mind.

I'm not the first one here. LiteLLM's proxy already does budgets and rate limits per key and per team, and most agent frameworks ship with iteration caps. Whatever you use, the advice holds: limit per agent, catch repeats, keep the meter outside the agent, and let the monthly cap be the backstop.

I'm Ken, a retired Navy Chief and 787 instructor. In aviation we don't trust one safeguard, we layer them. Same idea here.

If you want to follow along: https://flamingape.ai

How does your team handle runaway agents? I'd like to hear about it.

Drafted with AI help, reviewed and edited by me.

Top comments (0)