We launched a public overspend challenge on 2026-10-07: find a way to get an AI agent past the spending rules enforced by Pink Agentic AI Payments, our MCP server + REST API + rules engine for giving AI agents controlled spending power. We added a $100 bounty per verified win (first 3 wins, $300 total) on 2026-10-08.
On 2026-10-10 at 08:50 PT, two days after the bounty went up, GitHub user @ins0x4nur4g opened issue #1. No AI agent, no prompt injection. Plain curl against our REST API.
What happened
Our rules engine has a sensible policy: payments to known vendors under $1,000 auto-allow, and anything from $1,000 to $5,000 needs CFO approval. @ins0x4nur4g sent a payment request for an approved EUR vendor with the currency field written as "EUR ", a trailing space.
The REST edge didn't recognize the padded code, so it fell back to pricing the amount 1:1 as USD instead of converting it. A EUR 999 payment (about USD 1,079 at the engine's own EUR rate) got read as "USD 999", which is under the $1,000 auto-allow line. It skipped CFO approval entirely and got issued a single-use wire credential, test money, in our sandbox. They reproduced it on a second agent to confirm it wasn't a fluke, and kept the exact trigger out of the public issue, sending the details to us privately instead.
Roughly:
// request
{ "amount": 999, "currency": "EUR " }
// before the fix
// treated as USD 999 -> under $1,000 -> auto-allowed, no CFO hold
// after the fix
// HTTP 400: unsupported currency 'EUR '; use one of USD, EUR, GBP, HKD, SGD, JPY
The important detail: the MCP edge, which is what AI agents actually talk to, was never affected. It already enforced a strict currency allowlist. The bug was that our two entry points, REST and MCP, validated input differently. The rules engine itself made the right call every time; it just never got a chance to see the real currency.
Why it slipped through
Spending-policy engines get good testing on the "happy path" values builders expect: USD, EUR, exact matches. Whitespace, casing, and other near-miss inputs are the kind of thing that doesn't show up until someone deliberately goes looking, which is the entire point of running a public bounty instead of just internal QA.
The fix
We reproduced it the same morning and shipped a fix by about 10:00 PT, roughly 70 minutes after the report:
- Both REST endpoints now return HTTP 400 for any currency that isn't exactly
USD,EUR,GBP,HKD,SGD, orJPY. - The engine refuses to price an unknown currency code at all. It blocks, it never defaults to USD.
- REST and MCP now validate against one shared currency list instead of two separate ones.
We re-checked live after deploy: "EUR" correctly gets held for CFO approval, and "EUR ", " EUR", "EURO", and "XXX" all now return 400.
Same deploy, two smaller fixes worth mentioning: a reused idempotency key with a different payload now correctly returns 409 instead of silently replaying the original request, and daily/monthly spend counters now roll over automatically at the UTC day/month boundary instead of needing a manual reset.
The general lesson
This wasn't a rules-engine bug. The rules were correct. The input boundary in front of them wasn't. If you're building spending controls for an AI agent, or anything that gates money on a policy engine, three rules we'd pass on:
- Normalize and validate in exactly one place. If you have two entry points into the same policy engine, they will drift, and the drift is where the bug lives.
- Unknown input must fail closed. An unrecognized currency, vendor, or amount format should block the transaction, never fall back to a default unit or a permissive assumption.
- Assume the policy engine is tested more than the inputs to it. Attackers, and well-meaning bounty hunters, go looking for the places you didn't write a test for, not the rule you're proud of.
Pink Agentic AI Payments enforces spending rules on the server, in front of the money, so a fix to one rule boundary protects every agent at once.
Thank you, @ins0x4nur4g, for finding this cleanly, reproducing it, and keeping the trigger private until the fix was live. That's exactly how this is supposed to work. Issue #1 is now verified as win #1 of 3, and the hall of fame and fix log are updated in the repo.
Two bounties left, $100 each. If you can get an AI agent, or a plain HTTP client, past our spending rules, the challenge is open.
Our public sandbox is test money only: https://agentic-sandbox.pinkwallet.com. Production isn't live yet; we're running an early-access waitlist.
Update, same day: win #2 (the clock)
A few hours after this post went up, the same researcher filed a second report (issue #2). One of our sample rule sets has the rule "Outside 06:00 to 23:00 SGT: finance on-call decides". The server was checking that window against its own UTC clock, not Singapore time. At 02:07 Singapore time, a $1,000 supplier reorder that should have waited for the finance on-call person came back as allowed.
The report also pointed out a second piece of the same problem: on a real payment, the agent could send a local_hour value and pick which hour the rule saw.
Both were fixed the same day:
- Every workspace now has a timezone. It's shown on the workspace, in
GET /v1/rulesand in every decision trace. - Time-of-day rules run on the server clock in that timezone.
-
local_houronly works on dry runs (/v1/payments/check,pink.check_policy), to simulate an hour. A real payment ignores it and the trace says so.
Same lesson as the space character: the rule logic was fine, but the input it trusted (this time, the clock) wasn't. That's 2 of 3 bounties claimed. One $100 bounty is still open: overspend-challenge.
Top comments (1)
tr.ee/dev-to