On October 3, Simon Willison posted a one-liner that hit 400+ points on Hacker News: "We're going to need default hard budget caps on pretty much everything."
He's right. And if you run coding agents or deploy anything an agent built for you, this is your problem now.
Soft caps are a notification, not a control
"After $X, email me" is not a budget. It's a polite note that you're already broke.
A hard cap means: after $X this month, the service stops and returns errors. Willison's ask is simple. Most people would rather have their side project go down than wake up to a bill with an extra several hundred or several thousand dollars on it. If you want to remove the cap, you tick a box that says, roughly: "My app won't be shut down if I exceed the budget, and I'm responsible for the charges."
Opt-out, not opt-in. That's the whole proposal.
Why agents turn this from annoying to dangerous
Humans overspend slowly. Agents overspend at machine speed, and they never get bored of a retry loop.
Incidents people are circulating (these come from secondary write-ups, so treat the details accordingly):
- An autonomous agent scanning a hobbyist network reportedly ran up a $6,531 AWS bill by spawning duplicate CloudFormation stacks every time it hit an error. (source)
- Two agents, an Analyzer and a Verifier, reportedly ping-ponged requests for 11 days and generated a $47,000 bill. Each one looked fine from its own side. (source)
- One developer's refactor task normally cost ~180K tokens (~$1.40). On ambiguous files it hit 2.1M tokens before they killed it by hand. That's a 12x multiplier from a single confused loop.
The pattern is always the same: error, retry, error, retry, nobody watching. Now add that agents can provision paid infrastructure, not just call APIs. The thing that deploys the service also forgets the kill switch.
The clouds are moving, slowly
Credit where due:
- Google Cloud shipped Spend Caps in July: a monthly cap on specific services within a project.
- AWS launched project spend limits on September 16, as part of its new builder experience.
The AWS version is real, but read the fine print (breakdown):
| Detail | Reality |
|---|---|
| Default | Off. You configure it manually. |
| Scope | Max 10 projects per account |
| Rollout | Part of the new builder flow, not retrofitted to every existing account |
| Enforcement | Starts blocking new resources ~7 days before projected exhaustion; pauses EC2/RDS/Lambda/Bedrock/SageMaker ~4 days out |
| Paused projects | Data may be deleted after 90 days paused |
That's a genuine step forward and still not Willison's ask. The default still favors risk. If you're a new builder who never opens the settings page, you're back to hoping.
Stop waiting for vendors. Cap it in code.
Whatever your cloud does, your agent loop needs its own ceiling. Two rules:
- Enforce in code, never in the prompt. "Please stay under $5" in a system prompt is a suggestion to a model that's motivated to finish the task. It will route around it.
- Check before spending, not after.
Here's a minimal version. A budget that throws before the call, plus a velocity breaker to catch loops that stay under the total but burn hot:
class BudgetExceeded extends Error {}
class CircuitOpen extends Error {}
class TokenBudget {
private spent = 0;
constructor(
private maxUsd: number,
private usdPerMTokIn = 3,
private usdPerMTokOut = 15,
) {}
// call BEFORE every model/tool call
checkBefore(estimatedUsd = 0) {
if (this.spent + estimatedUsd > this.maxUsd) {
throw new BudgetExceeded(
`cap $${this.maxUsd} hit (spent $${this.spent.toFixed(2)})`
);
}
}
// call AFTER, with the usage object from the response
record(u: { input: number; output: number }) {
this.spent +=
(u.input / 1e6) * this.usdPerMTokIn +
(u.output / 1e6) * this.usdPerMTokOut;
}
}
class VelocityBreaker {
private window: number[] = [];
constructor(private maxTokPerMin = 10_000, private strikes = 3) {}
tick(tokensLastMinute: number) {
this.window.push(tokensLastMinute);
this.window = this.window.slice(-this.strikes);
if (
this.window.length === this.strikes &&
this.window.every((t) => t > this.maxTokPerMin)
) {
throw new CircuitOpen("sustained burn rate, likely a loop");
}
}
}
Plug checkBefore into the top of your agent loop and let the exception kill the run. Put your model prices in config, since they change.
Sane starting ceilings, adapted from the guard-pattern write-up above:
- Personal projects: ~50K tokens per session, ~$5/day
- Production SaaS: ~500K tokens per session, ~$100/day
- Batch pipelines: ~1M tokens per job, ~$500/day
Tune them, but have a number.
Checklist before you let an agent near a credit card
- [ ] Hard spend limit on every cloud project and every LLM provider key
- [ ] Per-run token ceiling enforced in code
- [ ] Max iteration count on every loop
- [ ] Alerts at 50/80/100%, but treat them as the backup, not the control
- [ ] Separate API keys per agent so you can revoke one without a fire drill
- [ ] Anything an agent deploys gets a cap at creation time, not "later"
The take
The next generation of builders will have agents create their infrastructure for them. If the default is "unlimited liability", some of them will get burned badly, and the first thing they'll learn about cloud is fear. Willison's point about AWS holds: people already refuse to use it for personal projects because of this.
Vendors should make caps the default. Until they do, you are the circuit breaker, and the safest place to put it is the line of code that spends the money.
What's the worst surprise bill you've seen? Drop it in the comments.
Top comments (0)