We gave an autonomous coding agent a real ticket: an off-by-one error in a pagination function, the kind of bug a mid-level engineer finds in about five minutes by reading the loop condition once. The agent took the ticket, read the codebase, formed a hypothesis, wrote a fix, ran the test suite, watched it fail, formed a new hypothesis, wrote a different fix, ran the suite again. Forty-five minutes and several dozen tool calls later, it had the right fix. It also had a bill that made someone in finance ask a question in Slack.
The wrong assumption going in
Our working assumption, based on the flat-rate subscription tooling we'd used before, was that agent time was basically free at the margin — slower than a human for genuinely novel problems, sure, but essentially costless to let run for a while on something simple. That assumption held for subscription-based tools with a fixed monthly price. It does not hold for pay-per-token autonomous agents working against a real API bill, where every tool call, every file read, every re-run of the test suite is metered.
Where the money actually went
The bug itself was genuinely small. The cost wasn't from solving a hard problem — it was from the agent's process for solving an easy one. Three things stacked:
It re-read far more context than it needed, repeatedly. Each tool call re-included substantial surrounding context, and the agent called tools far more times than a human would have needed steps, because it was exploring rather than reasoning from a mental model the way an experienced engineer would.
It ran the full test suite after every attempted fix, not just the relevant tests. A human fixing this bug runs the three tests that touch pagination. The agent, lacking the judgment call of "which tests are actually relevant here," ran everything, every time, because running everything is the safe default when you don't have confident judgment about what's relevant.
It didn't know when to stop and ask. After the second failed attempt, a human engineer would likely have paused and said "wait, let me actually read this function instead of guessing." The agent kept iterating through its hypothesis-fix-test loop, because that loop is what it's built to do, and nothing in its setup told it that the cost of the fifth iteration was meaningfully different from the cost of the first.
The fix
1. Hard budget ceilings per task, enforced before the agent starts, not reconciled after the invoice arrives.
MAX_TOOL_CALLS_PER_TASK = 15
MAX_COST_USD_PER_TASK = 2.00
def run_agent_task(task, budget_tracker):
while not task.is_complete():
if budget_tracker.tool_calls >= MAX_TOOL_CALLS_PER_TASK:
escalate_to_human(task, reason="tool call budget exceeded")
return
if budget_tracker.cost_usd >= MAX_COST_USD_PER_TASK:
escalate_to_human(task, reason="cost budget exceeded")
return
result = agent.step(task)
budget_tracker.record(result)
A $2 ceiling on a pagination off-by-one bug is generous. Hitting it is a signal to hand the task to a human, not a reason to raise the ceiling and let it keep going.
2. Scoped test runs instead of full-suite runs, inferred from which files the agent actually touched — cuts the compute cost of the verification loop dramatically, since most of the cost wasn't the fix, it was re-verifying the fix against everything every single time.
3. We moved agent execution into short-lived, disposable sandboxes sized specifically for the task, rather than a shared always-on environment billed by the hour regardless of whether an agent is actively working. I'm the founder of Krova Cloud, and this turned out to matter more than I expected for cost control specifically: a sandbox that exists only for the lifetime of one task, billed per minute, and destroyed the moment the task completes or hits its budget ceiling, means the infrastructure cost scales with actual agent activity instead of running as flat overhead whether the agent is doing five tool calls or five hundred. Combined with the budget ceiling, the pagination-bug-shaped task that used to cost real money now gets caught and escalated well before it burns through a meaningful fraction of what a human's five minutes would have cost in salary anyway.
Lessons
- Pay-per-token agent tooling changes the economics of "just let it try" in a way flat-rate subscription tooling never did. The habits that were harmless under one pricing model can be expensive under the other.
- An agent without an enforced budget ceiling will use exactly as many tool calls as its process calls for, with no awareness that the fifth attempt costs the same as the first. Build the ceiling in before the first real task, not after the first surprising invoice.
- Running the full test suite after every iteration is the safe default for an agent without judgment about relevance — and it's also often the single largest cost driver. Scoping verification to what actually changed is worth the engineering effort.
- The right response to an agent exceeding its budget on an easy task isn't a bigger budget. It's treating the overrun as a signal that the task should go to a human, the same way you'd treat a human engineer who'd been stuck on the same bug for three hours.
If you're running autonomous coding agents against a metered API without a hard per-task budget ceiling, it's worth finding out what your most expensive task looked like before it becomes a recurring line item.
I'm Rohit, founder of Krova Cloud — disposable, per-minute-billed sandboxes sized for exactly the task at hand, built so agent infrastructure cost scales with real activity instead of running as flat overhead. If you want more deep debugging stories like this one, I write regularly over at debugly.dev too.
Top comments (1)
The "ran the full suite every attempt" line is the one that stung — that was our single biggest cost leak too, and it's sneaky because it looks like diligence. What actually moved the number for us wasn't the dollar ceiling, it was giving the agent a changed-files → affected-tests map so it stopped running everything to re-confirm an unrelated green. The ceiling caught the runaways; test selection removed the baseline waste.
One thing I'd gently push on: a flat 15-call / $2 cap is great for an off-by-one, but we found a static ceiling either chokes a genuinely hard task or is too loose for an easy one. We ended up sizing the budget off a cheap up-front difficulty estimate, so the escalation trigger means "this is harder than it looked" rather than "you hit an arbitrary wall."
Curious — is your $2 a per-task constant, or do you scale it? And does hitting the ceiling route to a human, or to a cheaper model for a second opinion first?