DEV Community

vittoria
vittoria

Posted on

What happens when an LLM loop runs away: the guardrail pattern

Nobody budgets for the runaway loop. Every AI SaaS has a line item for "expected LLM spend" and nobody has a line item for "the Friday night a bug turned our agent into a money printer." I've seen the second one. Here's the pattern that prevents it.

The scenario

You ship a "deep research" agent endpoint on Friday at 6pm. It works like this: the agent plans, calls tools, reads results, and loops until its verifier step returns done: true.

Saturday morning, a customer pastes a URL the fetcher tool can't parse. The verifier keeps returning done: false because the evidence field is empty. The loop has no iteration cap — you meant to add one, the PR was getting long. The tool has retry logic with no backoff. The agent keeps going.

Nobody notices until Monday.

The math

Let's price one iteration on GPT-4o-class pricing ($2.50 / 1M input tokens, $10.00 / 1M output):

  • ~4,000 input tokens (growing context window) → $0.010
  • ~800 output tokens → $0.008
  • Per iteration: ~$0.018

At one iteration every 2 seconds, that's $0.009/second, or $32.40/hour — for a single stuck run. From Friday 6pm to Monday 10am is 64 hours:

64 × $32.40 = $2,073.60. From one user, one bad URL.

And that's the cheap version. If the trigger is a webhook redelivery storm instead of one user — say a provider retries a failed webhook 50 times and each delivery spawns an agent run — multiply accordingly. The failure mode isn't "slightly over budget." It's unbounded.

Why it always compounds

Runaway spend is never one bug. It's three missing defenses compounding:

  1. No iteration cap on the loop. The agent is the only thing that knows it's stuck, and nothing asks it to stop.
  2. Retry without a budget. Retries are priced in latency in every tutorial. In LLM systems they're priced in dollars.
  3. Metering as an afterthought. Usage gets logged somewhere — an events table nobody queries in real time. By the time the daily rollup runs, the money is gone.

The fix is a guardrail layer that sits between your code and the provider, and it has four parts.

The pattern

1. Price every call before you make it

Maintain a pricing table per model, versioned in code, updated when providers change prices. The critical rule: an unknown model must never silently price at $0. Either refuse the call or price it at the most expensive known rate. Silent $0 pricing is how a model-name typo becomes free unlimited usage — in the wrong direction.

PRICING = {
    "gpt-4o": {"input": 2.50, "output": 10.00},       # per 1M tokens
    "claude-sonnet-4": {"input": 3.00, "output": 15.00},
}
UNKNOWN_MODEL_POLICY = "refuse"  # or "price_at_max" — never "price_at_zero"

def estimate_cost(model, input_tokens, output_tokens):
    if model not in PRICING:
        if UNKNOWN_MODEL_POLICY == "refuse":
            raise UnknownModelError(f"No pricing for {model}; refusing call")
        rate = max(p["output"] for p in PRICING.values())
        return rate * (input_tokens + output_tokens) / 1e6
    p = PRICING[model]
    return (input_tokens * p["input"] + output_tokens * p["output"]) / 1e6
Enter fullscreen mode Exit fullscreen mode

2. Keep a pre-aggregated spend ledger

Do not SUM() the usage events table on every request. At any real volume that query is slow, and slow guardrails get skipped "temporarily" — permanently. Instead, maintain one row per API key per day (and per month): daily_spend, monthly_spend, updated on the metering write path after each call completes. The check before each call is then a single indexed row lookup — O(1), a few milliseconds, no excuse to skip it.

3. Check before the call, in middleware

The enforcement point belongs in one place — a FastAPI dependency or middleware — not sprinkled across every agent implementation:

async def guardrail_check(api_key, est_cost, db, redis):
    if await redis.get("llm:kill_switch"):
        raise HTTPException(402, "LLM spend globally paused")

    ledger = await get_spend_ledger(db, api_key.id)  # O(1) row lookup
    if ledger.daily_spend + est_cost > api_key.daily_cap:
        if api_key.enforcement == "hard":
            raise HTTPException(402, "Daily LLM budget exceeded")
        logger.warning("budget exceeded, allowing in warn mode",
                       extra={"key_id": api_key.id})

    if ledger.monthly_spend + est_cost > api_key.monthly_cap:
        raise HTTPException(402, "Monthly LLM budget exceeded")
Enter fullscreen mode Exit fullscreen mode

Two enforcement modes matter. Hard cutoff (HTTP 402) is what you want in production: the call never happens, the spend never occurs. Warn mode (log + allow) is what you want during development and for the first week after onboarding a big customer, so you can calibrate caps against real usage instead of guessing. Make it per-key, not global — your own internal tooling and your customers should not share a blast radius.

4. The kill switch

A single Redis flag — llm:kill_switch — that every enforcement point checks first. When the 3am page says spend is spiking and you don't yet know why, you flip one flag and all LLM spend stops in seconds, without a deploy. Set it, investigate, unset it. This is the cheapest incident-response tool you will ever build: about five lines of code, and it's the difference between a $200 incident and a $2,000 one.

Set caps with intent, too. Sensible starting points: a per-key daily cap around 3–5× the key's observed p99 daily spend, a monthly cap around 1.5× expected, and dedupe your threshold alerts (one Slack message per key per day at 80%, not one per request — alert fatigue is how the real warning gets missed).

One more calibration detail worth getting right: seed every new key's caps from the customer's stated expected usage, then auto-tune after the first week of real traffic. Caps set from guesses are either so high they never fire (useless) or so low they 402 legitimate traffic on day two (worse than useless — now support is involved). The warn-mode week exists precisely so you can watch real spend, set the daily cap at ~4× observed p99, and flip to hard enforcement with confidence.

The economics

Building this properly takes an afternoon: the pricing table, the ledger migration, the middleware, the kill switch, and tests that prove the cutoff actually fires. The incident it prevents costs $2,000 and a very bad Monday. That is the entire ROI argument, and I've never seen it lose.

I got tired of rebuilding this layer for every project, so I packaged it — pricing tables for OpenAI/Anthropic, the ledger, the middleware, hard/soft modes, the kill switch — into a FastAPI starter I sell as ShipSafe. But whether you buy mine or build your own this weekend, build it before the Friday deploy, not after the Monday invoice.

Top comments (0)