DEV Community

Quinn
Quinn

Posted on Originally published at pinkwallet.com

How to Stop an AI Agent From Overspending: 6 Failure Modes and the Controls That Catch Them

Disclosure: I work on Pink Agentic AI Payments at PinkWallet — mentioned once, in early access, near the end. Every vendor claim below links to that vendor's own page, checked on the date noted.

An AI agent overspends the same way a junior employee with a company card overspends: not through one dramatic fraud, but through a specific, boring failure that nobody built a check for. A cap set too high. A retry nobody deduplicated. A payee nobody validated. Below are six failure modes that actually cause agent overspend, a scenario for each, and the specific control that catches it — plus where that control needs to live to actually hold.

If you want the runnable code for a basic cap, see our earlier post, How to Give an AI Agent a Spending Limit; this one is about what a single cap doesn't catch.

1. Runaway loops: the agent keeps calling the same tool

Scenario: A research agent is told to "keep checking the API until you get a good result." A malformed response, a flaky endpoint, or a confused retry condition turns "keep checking" into hundreds of calls in a few minutes — and if each call is a paid API request, that's hundreds of small charges before anyone notices the agent is stuck.

Control: A velocity limit — a cap on the rate of spend or of tool calls (transactions per minute/hour), independent of any per-transaction amount. SpendNod's rule engine, for example, checks agent purchases against "per-transaction limits, daily/monthly caps, vendor blocklists, category restrictions, velocity limits" as one combined policy (spendnod.com, checked 2026-09-29) — velocity sits alongside the amount-based checks, not instead of them, because a runaway loop of small legitimate-looking charges won't trip a per-transaction cap at all.

2. Retries causing double charges

Scenario: The agent calls pay(), the network times out before the response comes back, and the agent (or your own retry logic) calls pay() again for the same intent — now there are two charges for one purchase.

Control: An idempotency key, checked before the payment executes, not just recorded after. Stripe's API docs describe the mechanism plainly: "The API supports idempotency for safely retrying requests without accidentally performing the same operation twice" — you attach a key to the request, and "if a connection error occurs, you can safely repeat the request without risk of creating a second object or performing the update twice" (Stripe idempotent requests, checked 2026-09-29). The catch for agent systems specifically: the payment API's own idempotency key doesn't help if your policy layer calls the API again with a new key because it didn't recognize the retry as a duplicate — the dedup has to happen at the policy layer too, before the call reaches the API.

3. Prompt injection steering payments to a new payee

Scenario: The agent reads a tool result, a scraped page, or an email — and buried in that text is an instruction to send a payment to a new "vendor." The agent's system prompt says "only pay approved vendors," but the injected text is competing for the same attention as that instruction, in the same context window.

Control: A merchant/destination allowlist enforced in code the injected text can never reach — not a rule the model has to remember and obey, but a check the payment call physically cannot bypass. This is a recurring pattern across vendor wallets: Circle Agent Wallets let you "configure allowlists and blocklists for wallet and contract addresses" (developers.circle.com/agent-stack/agent-wallets, checked 2026-09-29), and PayAgents lets you "construct clean domain track-record guardrails specifying exact allowed and disallowed destination domain lists" (payagents.io, checked 2026-09-29). An allowlist answers a question a cap can't: the amount can be exactly right and the payment still wrong, because it went to the wrong party.

4. Price or amount drift from what the user actually approved

Scenario: A user tells the agent "book the $40 flight," but by the time the agent executes the purchase, the fare has changed, or the agent picked a similar-looking item at a different price. Nothing here is malicious — it's just drift between what was approved and what's about to be charged.

Control: A per-transaction cap plus a lower approval threshold that routes anything above it to a human before it executes. Coinbase's Agentic Wallet documents "configurable caps per session and per transaction" (docs.cdp.coinbase.com/agentic-wallet/cli/welcome, checked 2026-09-29). SpendNod frames the same idea as three outcomes per transaction: "Auto-approved — Under your threshold — instant, no human needed," "Pending review — Above your threshold — you approve with one tap," or "Denied — Blocked vendor or limit hit — instant rejection" (spendnod.com, checked 2026-09-29). The cap is the hard ceiling; the threshold is what catches the "technically under the ceiling but still worth a human glance" case that drift produces.

5. A slow bleed of many small payments, each under the per-transaction cap

Scenario: Every single purchase is small enough to clear the per-transaction cap on its own — but the agent makes forty of them in a day, and the total is nowhere near what anyone approved.

Control: A rolling budget (daily/weekly/monthly), separate from the per-transaction cap. MPP's own guide to managing agent spend walks through exactly this: a scoped access key where "the key can spend up to 10 USDC per day and only transfer USDC to one recipient" (mpp.dev/guides/managing-agent-spend, checked 2026-09-29) — the daily figure is what stops the bleed, not the per-recipient restriction. Circle's Agent Wallets take the same approach at the rail level: spending limits "can be time-bound (for example, daily, or monthly)" (developers.circle.com/agent-stack/agent-wallets, checked 2026-09-29). PayAgents states it as a stacked set of windows: "you can enforce hard caps on a per-transaction, per-day, and per-30-days basis" (payagents.io, checked 2026-09-29) — three separate ceilings, because a per-transaction cap alone says nothing about total exposure over time.

6. No audit trail — you can't tell what happened after the fact

Scenario: Something goes wrong — a payment to an unexpected vendor, a cap that got hit, a spend spike — and the first question anyone asks is "which agent did this, under what policy, and was it approved?" If the answer isn't already recorded per attempt, you're reconstructing it from logs that weren't built for the question.

Control: A decision log for every payment attempt, not just the ones that succeeded — including denials and approvals-pending, since a wave of denied attempts is itself a signal something is wrong upstream. PayAgents describes this as "every transaction logged with proofs & hashes" (payagents.io, checked 2026-09-29). PolicyLayer takes a related but distinct angle — it's less a real-time gate and more "the system of record for AI agent authority," recording "your team's decisions, in one playbook every agent works from" (policylayer.com, checked 2026-09-29), so the org's spending rules and who decided them are themselves versioned and auditable, not just the transactions.

A minimal policy result already carries most of what an audit entry needs — you just have to actually persist it instead of discarding it once the payment either executes or gets denied:

// one line, appended per payment attempt — not just successful ones
auditLog.append({
  agentId, timestampUtc, amountUsd, merchant,
  decision: result.decision,      // allow | deny | require_approval
  reason: result.reason,
  policyVersion: policy.version,
});
Enter fullscreen mode Exit fullscreen mode

Where should the control actually live?

Not every layer that looks like a control actually enforces anything. Roughly three tiers, in increasing order of how hard they are to bypass:

  • The model / system prompt. "Don't spend more than $50" in a system message is not enforcement — it's a request the model can forget, misjudge, or have overridden by exactly the prompt-injection scenario in failure mode 3. Nothing here should be load-bearing.
  • The tool / MCP layer. A policy check that runs in code before the payment tool is allowed to fire — the pattern behind our own runnable example, and behind independent projects like agent-verifier-mcp, whose README frames it plainly: "MCP server for AI agent spending limits. One config change to add budget enforcement." Its Budget Authority Protocol is explicitly rail-agnostic: "Budget enforcement is rail-agnostic. Works with x402, credit cards, MPP, bank transfers — the verifier tracks the budget, the payment method doesn't matter" (github.com/goodmeta/agent-verifier-mcp, checked 2026-09-29). This is the layer that can enforce allowlists, approval thresholds and audit logging consistently across whatever rail a given payment happens to use.
  • The rail / wallet. Caps that live in the payment infrastructure itself, tied to the credential rather than to any code you wrote. These are real and they're the hard floor that holds even if your own policy code has a bug: Circle Agent Wallets, Coinbase's Agentic Wallet, and Tempo's wallet CLI all implement this independently. Tempo's docs state "each wallet can have multiple access keys with independent spending limits" (tempo.xyz/developers/docs/wallet/reference, checked 2026-09-29), and its CLI request flags include --max-spend, which will "stop if the request would exceed a spend cap," and --dry-run, which will "preview the payment cost and validate the request without spending" (tempo.xyz/developers/docs/cli/request, checked 2026-09-29). Coinbase and Circle both also screen transfers against OFAC/sanctions lists before submission onchain — Coinbase's docs state transfers are "automatically screened against OFAC sanctions lists and blocked before submission onchain," and Circle's state transfers "are screened against sanctions controls before submission onchain" (respectively: docs.cdp.coinbase.com/agentic-wallet/cli/welcome, developers.circle.com/agent-stack/agent-wallets, both checked 2026-09-29) — a control the tool layer can't replicate on its own.

Rail-level caps are genuinely useful and, where available, worth turning on regardless of what else you build — they hold even if your policy code ships a bug. Their limit is scope: a cap tied to one wallet or one rail's credential doesn't travel with the agent to a different rail, and it can't apply your organization's allowlist or approval-routing rules, which are business logic, not a property of the money. The case for putting the policy check at the MCP/tool layer is that it runs once, before the payment call, regardless of which rail ends up handling settlement — it's a complement to a rail-level cap, not a replacement for one.

Pink Agentic AI Payments (by PinkWallet, early access) runs three circuit breakers — monthly budget, company daily ceiling, vault balance — before any rule is even read, plus fraud checks (a payee's changed bank details, a duplicate invoice number) before amount tiers. Status: public sandbox live (test credentials only, no money moves); production not yet available.

Checklist

  • [ ] Per-transaction cap set to the price of the largest legitimate purchase, not a round number
  • [ ] Rolling (daily/weekly/monthly) budget, separate from the per-transaction cap
  • [ ] Merchant/destination allowlist checked in code, not just stated in the prompt
  • [ ] Approval threshold below the hard cap, routing mid-size or novel purchases to a human
  • [ ] Velocity limit on transaction rate, independent of amount
  • [ ] Idempotency key checked at the policy layer before the payment call, not just at the payment API
  • [ ] A decision log entry per payment attempt — denials and pending-approvals included, not just successes
  • [ ] Rail-level cap turned on wherever your payment provider offers one, as a second line of defense behind your own policy code

I work on Pink Agentic AI Payments at PinkWallet.

Related: Agentic AI Payments (2026): Protocols, Providers and Spending Controls · open guide: github.com/Pink-Agentic-Payments/agentic-ai-payments

Top comments (0)