DEV Community

Aman Kumar
Aman Kumar

Posted on

6 Guardrails That Stopped My Coding Agent From Burning Through API Budget

Coding agents are great until you look at the bill. A single "fix the failing tests" task can turn into 40+ model calls, each one re-sending the system prompt, file snippets and the previous tool output. Here are the six guardrails I now put on every agent loop, whichever model or provider sits behind it.

1. Cap iterations per task

Every agent framework has some notion of max steps. Set it explicitly. I use 20-30 tool turns for a normal task and stop the run if it is not converging. An agent that has not fixed a test in 25 turns is usually stuck in a loop, not about to succeed.

MAX_TURNS = 25
for turn in range(MAX_TURNS):
    step = agent.step()
    if step.done:
        break
else:
    log.warning("agent hit turn cap, stopping")
Enter fullscreen mode Exit fullscreen mode

2. Trim tool output before it goes back to the model

The biggest hidden cost is tool output. Do not send the whole test log or the whole file. Send the failing test names, the assertion lines and a diff.

def trim(output: str, limit: int = 4000) -> str:
    lines = [l for l in output.splitlines() if "FAIL" in l or "Error" in l or l.startswith(("+", "-"))]
    text = "\n".join(lines) or output
    return text[-limit:]
Enter fullscreen mode Exit fullscreen mode

3. Route by task type

Use a fast, cheap model for planning, search and summarising, and a frontier model only for the hard edit. With an OpenAI-compatible client this is just a different model string per call:

ROUTES = {"plan": "fast-model-id", "edit": "frontier-model-id", "review": "fast-model-id"}
client.chat.completions.create(model=ROUTES["edit"], messages=msgs)
Enter fullscreen mode Exit fullscreen mode

Get the exact ids from GET /v1/models on whatever endpoint you use; guessing ids is the most common setup failure.

4. Bound retries

Retry 429 and 5xx with exponential backoff, but only a few times, then fail over to a second model or stop. Unbounded retries inside an agent loop multiply spend without multiplying progress.

5. Watch the first real run

Before you leave an agent unattended overnight, watch one real task in the provider's request logs. You will spot runaway context growth, repeated identical calls and wrong model routing in minutes.

6. Pick a billing model you can forecast

Per-token pricing is fair but spiky for loops, because input tokens grow every turn. Two ways to make it predictable:

  • Budget alerts on your direct vendor account, with a hard monthly limit.
  • Flat-rate plans, where you pay a fixed monthly fee for a daily request allowance. You are still capped, but the bill does not move with context size. One example is APIClaw, an OpenAI-compatible gateway with flat monthly plans and one key across several model families (disclosure: I build it; it is independent and not affiliated with OpenAI or Anthropic).

If your usage is light or you need a first-party vendor contract, stay on the direct API and use budget alerts.

Checklist

  • Turn cap set
  • Tool output trimmed
  • Cheap model for planning, frontier model for edits
  • Bounded retries with a fallback
  • One supervised run checked in the logs
  • A bill you can predict

What guardrails do you use on your agents? I would love to hear what I am missing.

Top comments (0)