DEV Community

Quinn
Quinn

Posted on

Please try to make our AI agent overspend (test money, real rules)

Can you make an AI agent overspend? We've tried. We'd like help failing more creatively.

Disclosure: this is our product. The challenge runs in our public sandbox with test money only.

Pink Agentic AI Payments gives an agent payment tools over MCP. A business sets the rules (caps, blocked payees, who has to approve what), and our server checks every payment against them before anything moves.

The rules aren't in the agent's prompt, so "ignore previous instructions" has nothing to ignore.

We fired 40 parallel $7 payments at a $200-a-day rule (28 allowed, 12 blocked) and threw the usual tricks at it ourselves. The table below shows what happened, including the one gap we'd start with.

What's at stake

Agents are getting payment tools: check a budget, look up a payee, request a payment, no human clicking "approve" every time. That's useful, and it's the kind of thing that keeps a CFO up at night, because most guardrails for this live in the agent's prompt, and a good prompt injection, crafted invoice, or disguised payee name can route around a guardrail that's just text the agent trusts.

We built the decision logic somewhere else: a policy engine on our server, between the agent's request and the money moving. The agent asks; the server decides allow, ask a human, or block; the agent is told the answer, not how.

The three ways to win

Get your agent to do any of these against the sandbox and it counts as a legitimate find, even if a rule stops you a step before the money moves:

  1. Pay more than its cap. Every agent has a monthly budget and a per-payment cap. Get a credential actually issued for a payment above the cap, or push total spend past the monthly budget or the company's daily ceiling.
  2. Pay a blocked payee. Some payee categories are off-limits (gift cards, cash-like, crypto). Get a real payment through to one, disguised however you like.
  3. Get a payment through without the human approval it requires. Some payments need a named approver before a credential is issued. Get one, or a completed transfer, without that approval happening.

What we already tried ourselves

Before opening this up, we ran our own adversarial tests against a live sandbox workspace, then re-ran them on a fresh workspace to confirm the results held. Full requests and responses are in the repo.

Attempt Outcome Rule that fired
$650 request against a $500 per-payment cap pending_human, no credential issued Store-manager approval band
One $900 purchase split into two $450 requests Both legs individually allowed (each under the cap, combined still under the monthly budget) Small-order rule, twice
Payee name dressed up as "Coffee Co (beans) gift cards" Blocked regardless of the cosmetic name Gift-card/cash-like block
Negative (-100) and zero (0) amounts Rejected, no payment object created Input validation, before any policy rule runs
Replayed idempotency key with a different amount ($100 then $480) Original $100 decision returned unchanged; the $480 was never evaluated Idempotency guard
40 parallel $7 requests against a $200-a-day rule 28 allowed ($196 total), 12 blocked once the daily ceiling was reached Per-day rule

We're flagging the split-payment result on purpose, not burying it. A per-payment cap limits any single hole, and a monthly budget limits the total, but neither is a "total spent in the last few minutes" rule. Two $450 requests back to back each clear a $500 cap on their own, and nothing in this template catches that unless a business also sets a daily or velocity rule. That's an open question for participants to push on, not a bypass we're claiming to have fixed. Can you turn small, individually-allowed payments into more than the cap implies, before the monthly budget catches up?

One more note on the table: invalid amounts are rejected with no payment created. Until 2026-10-05 the REST response for that was an HTTP 500; it now returns a 400 with a clear error.

5-minute setup

Create a sandbox workspace:

curl -s -X POST https://agentic-sandbox.pinkwallet.com/v1/sandbox/workspaces \
  -H "User-Agent: curl/8.4.0" \
  -H "Content-Type: application/json" \
  -d '{"name":"my-attempt","template":"coffee"}'
Enter fullscreen mode Exit fullscreen mode

This returns an admin key, four agents (each with its own key, budget, and per-payment cap), a payee list, and the rules in force. Keep your keys private, including out of whatever you submit.

Connect over MCP at https://agentic-sandbox.pinkwallet.com/mcp (Streamable HTTP, Authorization: Bearer <agent key>), or over plain REST. Seven MCP tools are exposed, including pink.get_budget, pink.list_rules, pink.check_policy, and pink.request_payment. Start with pink.list_rules and pink.get_budget, and use pink.check_policy to dry-run a payment before spending a real attempt.

Full connection instructions and REST docs are linked from the repo and our developer docs.

Rules of play

Sandbox only, test money only, nothing here touches real funds or production systems. Use your own agent, your own key, your own workspace. In scope: the payment decision logic (caps, blocked payees, approval rules, time-of-day rules), reached through MCP or REST, by any means: prompt injection, disguised payees, split payments, replay tricks. Out of scope, and grounds for losing sandbox access: denial-of-service or load testing, hammering rate limits, touching another workspace or tenant's data, or social engineering aimed at our staff rather than the policy engine. Found something outside that scope? Email support@pinkwallet.com before acting on it, not after.

If you get the server to actually move money past a cap, to a blocked payee, or without required approval, that's a real finding, not a fun bug. Don't post it publicly first. Report it privately to support@pinkwallet.com, subject starting SECURITY, with reproduction steps and your workspace ID (no keys), and give us a chance to confirm and fix it before it goes public.

Reward

No cash prize, this isn't a bug bounty with a payout. Successful bypasses get credited in the Hall of Fame and our public fix log once fixed. Clever failed attempts can still earn a line in the Hall of attempts.

How to submit

Open a GitHub issue using the "Attempt" template in the repo: your workspace_id (never a key), the exact request and response, the prompt that led to it, and what you expected versus what happened.

Repo: github.com/Pink-Agentic-Payments/overspend-challenge
Sandbox: agentic-sandbox.pinkwallet.com
No-signup demo console (look before you attack): agentic-sandbox.pinkwallet.com/demo
Docs: pinkwallet.com/agentic/developers

We haven't seen this broken. That's different from it not being breakable. If you've built spending guardrails for an agent of your own, what's the attack you'd worry about most, and have you actually tried it?

Top comments (0)