In traditional software, a badly written loop gives you a CPU spike. Somebody gets paged, the graph goes back down, and the cost is a slow afternoon.
In an AI system, the same mistake is a billing event. Every iteration is a paid API call, and the meter doesn't care whether the call was useful. A loop that runs for ten minutes before anyone notices can cost more than the feature earns in a month.
The tutorial version of an LLM feature has one call, one response, done. The production version has retries, self-correction steps, agents calling tools that call models, and thousands of users doing all of that at once. Every one of those multipliers is a place where cost can run away from you.
Here are the three failure modes I see most often.
Failure mode 1: the self-correcting loop that never converges
Illustrative scenario: an agent generates JSON, validates it, and if validation fails, sends the error back to the model and asks it to fix the output. Reasonable pattern. Then a schema change ships, and one field becomes impossible to satisfy. The model keeps "fixing" the output, the validator keeps rejecting it, and the loop keeps going.
Nothing crashes. No exception reaches your error tracker. The only signal is the invoice.
Fix: every retry and correction loop gets a hard ceiling, in attempts and in tokens, and hitting the ceiling is a logged failure, not a silent fallback.
async function generateValid<T>(prompt: string, validate: (x: unknown) => T) {
const MAX_ATTEMPTS = 3;
const MAX_TOKENS_TOTAL = 20_000;
let tokensUsed = 0;
let lastError = "";
for (let attempt = 1; attempt <= MAX_ATTEMPTS; attempt++) {
const res = await callModel(prompt, lastError);
tokensUsed += res.usage.totalTokens;
if (tokensUsed > MAX_TOKENS_TOTAL) throw new BudgetExceeded(tokensUsed);
try {
return validate(res.output);
} catch (e) {
lastError = String(e);
}
}
throw new MaxAttemptsExceeded(MAX_ATTEMPTS, lastError);
}
The same applies to agents: cap the number of tool calls and model calls per task, not just per request.
Failure mode 2: traffic scales linearly, cost doesn't
A traffic spike on a normal endpoint costs you some extra compute. A traffic spike on an endpoint that fans out to several model calls per request, each with a long context window, costs you multiples of that.
If one user action triggers a summarization call, a classification call, and an embedding call, your cost per request is the sum of three metered APIs. Double the users and you've doubled all three, before retries.
Fix: know your cost per request the same way you know your latency per request. Put a budget on the feature, and decide in advance what degrades when the budget is hit: a cheaper model, a shorter context, a cached answer, or no AI output at all.
Failure mode 3: no limits per user
Global rate limits protect your provider quota. They don't protect you from one user (or one script, or one bug in your own frontend that re-fires a request on every render) consuming a disproportionate share of it.
Illustrative scenario: a chat feature with no per-user cap. A single integration test, pointed at production by mistake, sends requests in a tight loop over a weekend. The global limit is never reached, so nothing alerts.
Fix: rate limits and token quotas per user, per API key, and per tenant. Treat them like you treat auth: not optional, and enforced at the edge, not deep in a handler.
Monitoring is now a finance problem
Uptime dashboards tell you whether the system is running. They don't tell you what it cost to keep it running in the last hour. For AI features, you need both:
- tokens and spend per feature, per user, per tenant, in near real time
- alerts on spend rate, not just error rate
- a kill switch that turns a feature off when it crosses a threshold, without a deploy
The reframe: an LLM call isn't a function call. It's a purchase. Your architecture should treat every loop, retry, and fan-out as a spending decision with a limit attached.
If you don't have hard cost controls in production, you're not running a feature. You're running an open tab.
What's the most expensive loop you've seen an AI feature get stuck in, and what limit would have caught it?
Top comments (2)
The ceiling in generateValid is the right instinct, but it bounds one retry layer, and most stacks have at least two more.
Below it, the provider SDK. The OpenAI and Anthropic Python SDKs both default to max_retries = 2. Those only fire on connection errors, 429s and 5xx, so they mostly cost latency, but those are exactly the errors you get during the spike from failure mode 2.
Above it, whatever runs the task. If generateValid throws MaxAttemptsExceeded inside a background job, Sidekiq for one defaults to 25 retries with backoff. Your schema-change scenario becomes 3 attempts × 26 runs = 78 paid calls per task, spread over about three weeks, with every layer individually "capped". Spread that thin, it never shows up as a spike.
Two things I would add to the fix:
The budget belongs to the task, not to the function call. A counter created inside generateValid starts at zero every time the job re-runs. Store it with the job (or under a key with a TTL) and have every layer draw from it.
BudgetExceeded and MaxAttemptsExceeded should be non-retryable at the queue level: straight to the dead set, with an alert. Retrying a deterministic failure is the expensive way of doing nothing.
A small one on the snippet: the check runs after the call, so the call that crosses 20,000 is already paid. Capping each call's output at what is left of the budget, minus the prompt, makes the ceiling a real one.
Have you seen spend-rate alerts catch the slow version of this, or only the spikes?
Spot on with the background job multiplier. That Sidekiq interaction is terrifying because it turns a fast bug into a multi-week slow drain that completely bypasses rate spikes. Good catch on the code snippet too, since checking after the call means that last over budget payload is already charged.
To answer your question, spend-rate alerts usually just tell you when you are already bleeding, unless they trigger an automated circuit breaker. The slow burns tend to slip past alerts because they look like normal organic growth until the bill lands. Have you found any alerting tools that actually catch those slow leaks early, or do you just rely on hard daily caps?