DEV Community

The Agent Loop
The Agent Loop

Posted on

Where does the budget check go?

Drafted with AI help, human-reviewed by The Agent Loop.

Short version: Your spend limit lives at the month. Your bill is generated at the call. The check has to live where the call happens, and almost nothing there is asking.

The prompt that exists is not the budget

You already set a spend limit, so you already believe you have a budget check. Read what it actually does: OpenAI's hard limit sits at organization or project scope, monthly, and when it trips, "affected API requests return a 429 error with the organization_spend_limit_exceeded or project_spend_limit_exceeded code. Enforcement is not instantaneous, so recorded spend can slightly exceed the configured amount" (OpenAI docs, 2026-09). The limit is real. It is also a month wide.

Meanwhile the only prompt your agent sees before acting asks one question: may this tool run? MCP's security docs require explicit consent before a tool executes, and Claude Code can be configured to ask on every call (Anthropic docs, 2026-09). Read the specs: consent is allow/deny. We found no monetary cap field anywhere in the MCP authorization or security material (reviewed 2026-09). Your approval dialog for a one-cent call and a forty-dollar call is the same dialog. I have written before about approval gates that never fire; this is the harder version. A consent prompt can fail to appear. A budget check that was never specified cannot fail to appear. It was never there.

Four places a check can live

Before you can place the check, you have to say what kind of check you mean. There are only four slots:

  1. Pre-call reservation. The gate: check remaining budget, then let the model run. This one exists. OpenAI's Agents SDK blocking-mode guardrails "run and complete before the agent starts. If the guardrail tripwire is triggered, the agent never executes, preventing token consumption and tool execution. This is ideal for cost optimization" (OpenAI docs, 2026-09). Tool guardrails wrap the call itself: input check before execution, output check after.
  2. Post-call ledger. Record what was spent and reconcile it against what you expected. The payments world has this: x402 documents settlement reconciliation with an exactly-once retry (x402 docs, 2026-09). Cost ledgers for LLM calls at run scope are mostly still your own plumbing (see the attribution post).
  3. Retry budget. The silent multiplier. LangChain's AgentCore Payments middleware enforces "spending limits at the session level before any payment is signed", validates each payment against the session budget and rejects it when the limit is exceeded, and defaults max_error_retries to 3 per tool call with an auto-created session budget of one dollar (LangChain docs, 2026-09; preview feature). That default of 3 is exactly the shape your non-payment retries should have too.
  4. Subagent envelope. Give each fork its own remaining budget so N children cannot each inherit an open-ended parent. In the SDK docs we reviewed, no dedicated per-subagent dollar budget or fork-count primitive exists; the recommendation falls back to tool guardrails around delegated calls (OpenAI docs, 2026-09). Slot 4 is where you build your own.

MCP fills none of these at protocol level. What exists instead is the gateway pattern: per-user "buckets to contain a runaway agent", "cost-weighted per-tool buckets to protect expensive operations", per-tenant isolation (Scalekit, 2026-06). That is application code in front of the protocol, which is fine, as long as you know the protocol itself is not coming to help.

What each provider's "spend limit" actually does

Provider What actually fires What never happens
OpenAI 429 with *_spend_limit_exceeded at org/project monthly scope, slightly late (docs, 2026-09) Per-call dollar kill
Anthropic Tier spend cap pauses all API usage until 00:00 UTC on the first of next month (429); your own spend limit returns HTTP 400 (docs, 2026-09) Per-call dollar kill
Azure Notification: "Resources aren't affected, and your consumption isn't stopped." Budgets evaluate every 24 hours, against data that lags 8–24 hours (docs, 2026-06) Stopping anything
Google Cloud After 100% of budget, pauses new usage for one eligible service in one project (Gemini API, Vertex/Agent Platform, Cloud Run); in-flight requests finish (docs, 2026-09) In-flight termination; multi-service scope
AWS Budget actions apply an IAM policy or SCP, e.g. deny provisioning more EC2 (docs, 2026-09) Per-inference dollar cap

The strongest control in this table (GCP's) still lets every in-flight request finish, covers one service in one project, and waits for you to lift it manually. The weakest (Azure's) evaluates a data set that is already up to a day old and stops nothing. No provider documents a per-call monetary kill. The month is where they enforce; the call is where you pay.

The gap, with a receipt

A reported incident from earlier this year: an operator's AI agent set out to join and scan the DN42 network, and the AWS bill attached to the account came to $6,531.30 (lantian.pub, 2026-05; the operator's own account, not an audit). I use it carefully: it does not prove a missing per-call budget was the sole cause, and I am not aware of any audited postmortem that does. What the account does describe is inadequate stopping controls: the system that knew the spend kept going because nothing between the agent and the account could say "you have used your allowance for this task, stop."

Press reports of seven-figure agent bills circulate every month. None of them are audited postmortems. The documented controls tell the same story without the drama: enforcement at the month, consent at the tool, nothing at the call.

Four moves that put the check in the right slot

  1. Gate before the model, after the arithmetic. Put a blocking pre-agent check (guardrail mode or your own hook) that compares remaining run budget to an estimated next step, and refuses before tokens burn. This is the only slot where stopping is free.
  2. Number your retries. Cap retries per tool call at a small constant (three is the published default above) and count each retry against the run's budget. Retries are how a loop becomes a bill without a single decision looking expensive.
  3. Envelope every subagent. Slot 4 has no framework primitive, so pass an explicit remaining-budget value down to each fork, and refuse to spawn a child without one. A child that inherits an open-ended parent multiplies the parent.
  4. Keep the provider limit, and stop trusting it. It is the last-ditch 429 when your own gate fails, not the gate. If you are on Azure, your budget is an email; write your own check.

FAQ

Is this just rate limiting? Rate limits protect the provider from you. A budget check protects you from your agent: rate limits count requests per minute, budgets count money per task. You need both, and only one of them is priced.

Where does MCP fit? It standardizes consent and authorization scopes so tools can be gated. Spending ceilings are left to the authorization server and your host code, which means today they are your job.

Doesn't this slow the agent down? The pre-call check is one subtraction and a comparison against a number you already have. The Azure budget is evaluated every 24 hours. You are trading microseconds per call for protection between calls.

Sources


Where does your budget check live: before the call, after the call, or in the postmortem? Tell me which slot you actually have today.

Related: Why your MCP approval gate never fires · What one agent run actually costs · Your agent's cost problem isn't the model. It's the loop.

If this saved you from one unbounded loop, tap the unicorn below; it takes one click and it is the only metric Dev.to shows me. And follow The Agent Loop if you want tomorrow's post in your feed: I am working through where money actually moves in an agent, one post a day.

Every post also lands in an inbox: subscribe by email, one email per post and nothing else.

Top comments (0)