Disclosure: I work on distribution for PinkWallet, which is building an agentic payments product (mentioned once near the end, honestly, as one option in early access). Every third-party claim below links to that vendor's own docs, checked on the date noted.
Short answer: A spending limit only counts if it's enforced somewhere the agent can't talk its way around — which in practice means a policy check that runs in code, outside the model's context, before a payment tool call is allowed to execute. Telling the agent its budget in the system prompt is not a spending limit; it's a suggestion the model can forget, get talked out of, or have overridden by a prompt injection. Below: where a limit can actually live, the specific controls you need, a runnable policy-check wrapper you can paste and run, and how today's payment providers handle (or don't handle) this.
Where a limit can actually be enforced
There are five places a "don't spend more than X" rule can live. Only some of them are enforcement; the rest are advice.
| Layer | What it can enforce | Bypassable by the model? | Example |
|---|---|---|---|
| The system prompt | Nothing, technically — it's a request | Yes, trivially (bad instruction-following, prompt injection, or the model just being wrong) | "You may spend up to $50" in a system message |
| Agent framework / tool wrapper | A code-level check before the tool function runs | No, if the agent has no other path to the underlying payment call | A custom pay() wrapper like the one below; LangChain/LlamaIndex tool middleware you write yourself |
| MCP server / tool layer | Whatever the server's own tool implementation checks before calling the payment API | No, if the server is the only credentialed path — yes, if the agent also has direct API access | Stripe's MCP server requires human confirmation before certain stripe_api_write actions such as refunds and outbound payments (Stripe MCP docs) |
| Payment provider | Caps, allowlists, or scopes tied to the credential itself, independent of any code you write | No — this is the hard floor even if your own code has a bug | Circle Agent Wallets expose per-transaction/daily/weekly/monthly spending limits and address allow/blocklists (Circle, GitHub); Coinbase's Agentic Wallet documents configurable per-session and per-transaction caps (Coinbase docs) |
| Card network / bank rail | Whatever the issuing bank or network allows on that specific card or account | No | A virtual card with a hard limit set by the issuer |
The practical rule: the system prompt is not a control, it's documentation. Real enforcement starts at the tool-wrapper layer and gets stronger the closer it sits to the actual money movement. The MCP landscape is uneven here — most payment-capable MCP servers we checked ship write tools (refunds, payouts, order creation) that execute immediately once a credential is configured, with the spend limit left entirely to whatever scope you put on that credential, not to a purpose-built budget feature inside the MCP server itself. Stripe and Airwallex were the two exceptions we found with a built-in confirmation gate: Airwallex's agent tools "do not initiate money-out actions (transfers, FX conversions, or payouts)" by default and prompt for confirmation on write actions (Airwallex AgentOS docs).
That's why you generally want at least two of these layers active at once: a code-level check you control (fast to change, but only as strong as your code), plus a provider- or network-level cap (slower to change, but holds even if your code has a bug).
The policy primitives
These are the individual controls a real policy is made of. Mix and match based on what the agent actually needs to do — a research agent buying API credits and a procurement agent paying SaaS invoices don't need the same numbers.
- Per-transaction cap — the hard ceiling on any single payment. Set it to the price of the largest legitimate item the agent buys, not a round number.
- Rolling budget — a daily/weekly/monthly total, separate from the per-transaction cap, so an agent can't stay under the per-transaction limit while still spending unbounded amounts over time.
- Merchant/endpoint allowlist — constrains who gets paid, not just how much. A cap alone doesn't stop the right amount going to the wrong party.
- Category blocks — a second-layer blocklist for things that should never be payable regardless of amount (payroll, gift cards, crypto exchanges), because an allowlist alone can't anticipate every bad actor you'd want to exclude by category.
- Human-approval thresholds — not a hard block: it routes the transaction to a human for a yes/no before it executes. Lower than the per-transaction cap, so mid-size or novel purchases still get a look.
- Velocity limits — caps the rate of spending (transactions per hour/day), independent of size. A per-transaction cap doesn't stop 200 small, individually-legitimate-looking payments in an hour.
- Kill switch — an immediate way to stop an agent from transacting at all, separate from normal credential rotation.
- Audit log — a record, per transaction, of which agent acted, under what policy, for how much, to whom, and with what approval — the thing you need after something goes wrong.
Velocity limits and dedicated kill-switch workflows are the two controls we found least documented among payment providers as of this writing — most leave both to you to build at the policy layer. (Full source list in the spending policy template.)
A runnable example
This is a minimal, provider-agnostic policy check in plain Node.js — no SDK, no network calls. The pattern is what matters: the agent's "pay" tool is never the raw payment call, it's a wrapper that checks policy first and only calls the real function if the check passes.
policy.js — the check itself:
export function checkPolicy(policy, req, ledger) {
if (req.amountUsd <= 0) {
return { decision: "deny", reason: "non-positive amount" };
}
// A retried request with the same key must not be treated as a new charge.
if (ledger.alreadyProcessed(req.agentId, req.idempotencyKey)) {
return { decision: "deny", reason: "duplicate idempotency key — already processed today" };
}
const merchant = req.merchant.toLowerCase();
if (!policy.merchantAllowlist.includes(merchant)) {
return { decision: "deny", reason: `merchant "${merchant}" is not on the allowlist` };
}
if (req.amountUsd > policy.perTransactionCapUsd) {
return {
decision: "deny",
reason: `amount $${req.amountUsd} exceeds per-transaction cap $${policy.perTransactionCapUsd}`,
};
}
const projectedTotal = ledger.spentTodayUsd(req.agentId) + req.amountUsd;
if (projectedTotal > policy.dailyCapUsd) {
return {
decision: "deny",
reason: `would bring today's total to $${projectedTotal}, over daily cap $${policy.dailyCapUsd}`,
};
}
if (req.amountUsd >= policy.approvalThresholdUsd) {
return {
decision: "require_approval",
reason: `amount $${req.amountUsd} is >= approval threshold $${policy.approvalThresholdUsd}`,
};
}
return { decision: "allow", reason: "within cap, allowlisted, below approval threshold" };
}
pay-tool.js — the wrapper the agent actually calls. Note there's no code path from the agent to executePaymentUnsafe that skips checkPolicy:
export async function pay(policy, req, ledger, onApprovalNeeded) {
const result = checkPolicy(policy, req, ledger);
if (result.decision === "deny") {
return { status: "denied", reason: result.reason, request: req };
}
if (result.decision === "require_approval") {
await onApprovalNeeded(req);
return { status: "pending_approval", reason: result.reason, request: req };
}
const executed = await executePaymentUnsafe(req);
ledger.record(req.agentId, req.idempotencyKey, req.amountUsd);
return { status: "allowed", reason: result.reason, result: executed };
}
Running eight scenarios against one agent (per-transaction cap $50, daily cap $90, allowlist of three vendors, approval threshold $40) against this wrapper, on Node v22.21.0:
1. Normal purchase, allowlisted, under threshold: allowed — within cap, allowlisted, below approval threshold
2. Same amount again but new key (should still be allowed, under daily cap): allowed — within cap, allowlisted, below approval threshold
3. Retry with the SAME idempotency key as #1 (simulates a network-retry double-charge attempt): denied — duplicate idempotency key — already processed today
4. Amount above per-transaction cap: denied — amount $75 exceeds per-transaction cap $50
-> [approval channel] would notify a human about: $45 to aws.amazon.com
5. Amount at/above approval threshold but under per-transaction cap: pending_approval — amount $45 is >= approval threshold $40
6. Merchant not on the allowlist (simulates a prompt-injected redirect): denied — merchant "totally-not-a-scam.example" is not on the allowlist
7. Another normal purchase (brings today's total to $75 of the $90 daily cap): allowed — within cap, allowlisted, below approval threshold
8. Legitimate small purchase that would push the daily total over the cap: denied — would bring today's total to $105, over daily cap $90
Total spent today for research-agent-01: $75
That's the real output from running node demo.js locally. The full source (policy.js, pay-tool.js, demo.js) and a README are in agent-spending-limit-example (MIT). The SpendLedger here is an in-memory Map, which is fine for a demo and not fine for production — see the race-condition note below.
Common failure modes
- Prompt injection → spend. A tool result, a scraped webpage, or a document the agent reads can contain text instructing it to make a purchase. The system prompt telling the agent "only buy from the allowlist" doesn't stop this — the allowlist check has to live in code the injected text can't reach, which is exactly why the wrapper pattern above puts the check outside the model's context entirely.
-
Retries → double charges. Network timeouts and agent retry loops are common; if your "did this succeed?" check happens after a timeout, the agent (or your own retry logic) may call
payagain for the same intent. An idempotency key checked before execution (scenario 3 above) is the fix — first-class API-level idempotency keys are a common feature (see Stripe's idempotent requests), but you need your policy layer to honor them too, not just the payment API. - Currency confusion. A cap defined as "50" without a currency is a bug waiting to happen the first time an agent operates against a non-USD price or a stablecoin quoted in a different denomination. Every cap and log entry needs an explicit currency field.
-
Budget race conditions. If two payment requests for the same agent are checked concurrently, both can read the same "spent so far" total, both pass the daily-cap check, and both execute — busting the cap. The in-memory
Mapin the demo above is illustrative only; a real ledger needs an atomic increment-and-check (a database transaction, a single-writer queue, or a provider that enforces the cap itself as a second line of defense).
What today's providers actually let you enforce
Verified against each vendor's own docs (checked 2026-09-27/28; not re-verified since):
- Circle Agent Wallets document per-transaction/daily/weekly/monthly spending limits and address allow/blocklists, and gate limit changes behind an OTP the agent never sees directly ("the agent hands the user a verbatim command to run in their own terminal so the OTP never passes through agent storage") — github.com/circlefin/skills.
- Coinbase's Agentic Wallet documents configurable caps per session and per transaction — Coinbase docs.
- Stripe's MCP server requires human confirmation before certain write actions like refunds and outbound payments, via a link that expires after 24 hours if unconfirmed; starting October 31, 2026, it also stops accepting full-access secret keys or restricted keys without an Agent tag (Stripe MCP docs).
- Crossmint describes agent credentials as "scoped, explicit, and revocable," using one-time or encrypted credentials rather than a raw card number — Crossmint docs.
- Skyfire lets you set spending limits per agent at the wallet/dashboard level — Skyfire product page.
- x402 (the pay-per-request HTTP protocol) settles one payment per request and doesn't itself carry a budget concept — any per-agent cap has to be enforced by whatever signs the payment on the agent's behalf, one layer above the protocol (x402 FAQ).
- Most traditional processors we checked — Adyen, PayPal, Square, Checkout.com, Razorpay — inherit spend control from the merchant's existing API-key or dashboard-role permissions, which were designed for human developers, not for bounding one autonomous agent's spend. None of the pages we checked for those five described a per-agent budget feature.
Where this leaves you: provider-level caps (Circle, Coinbase) are the hard floor that holds even if your own code breaks, but they're not universal across vendors yet, and velocity limits and dedicated kill switches are thin on the ground industry-wide. That's the gap a code-level policy wrapper like the one above is for.
PinkWallet is building Pink Agentic AI Payment, an MCP server plus a console for defining exactly this kind of rule — per-agent caps, allowlists, and approval thresholds enforced at the MCP layer before a payment executes. It's in early access — not live yet, no claims otherwise — alongside the other options named above.
FAQ
Is putting the limit in the system prompt good enough for a low-stakes agent?
No, for the same reason input validation belongs in code, not in a comment telling users to behave. Even a "low-stakes" agent can be redirected by a prompt injection it reads from a tool result; a system-prompt limit has no mechanism to stop that.
Do I need both a provider-level cap and my own policy code?
Where you can get both, yes. The provider-level cap (Circle's or Coinbase's, for example) is the floor that holds even if your policy code has a bug; your own policy layer is what lets you apply consistent rules — allowlists, approval routing, velocity limits — across multiple providers and agents, which no single provider's cap will do for you.
What's the difference between a cap and an approval threshold?
A cap is a hard limit the agent cannot exceed. A threshold is lower and doesn't block anything — it routes the transaction to a human for a decision before it executes. Most real policies need both, at different amounts.
Can velocity limits be enforced by the payment provider instead of my own code?
Not reliably today — velocity limits were the least-documented control among the payment providers checked for this article. Plan to implement rate limiting for agent spend at your own policy/orchestration layer rather than assuming a provider handles it.
Does an MCP server automatically give me spending controls?
No. Most payment-capable MCP servers we checked ship write tools that execute immediately once a credential is configured; the spend limit comes from how narrowly you scope that credential and whether you've added your own policy check in front of the tool call, not from the MCP protocol itself.
Top comments (0)