DEV Community

Jiahui Miao
Jiahui Miao

Posted on

Stop Giving Your Agent a Wallet. Give It a Treasurer.

Stop Giving Your Agent a Wallet. Give It a Treasurer.

This week we shipped the least glamorous module in our entire stack, and the one I'd now refuse to run without: a treasurer.

Not a wallet. A treasurer. A small, boring, mechanical gate that sits in front of every external paid call my agent wants to make and asks a few dumb questions: is this payee on the whitelist? Is the amount under the per-transaction cap? Under the daily and monthly caps? Same recipient name but a different destination account than history shows? That's not a "no" — that's an escalation, straight to a human review queue with full context. And if the gate itself is misconfigured or confused, the answer is always no. Default-deny. Fail closed.

Twenty tests, all green. Wired at the head of the guardrail chain, so nothing paid can run before it clears. Conservative limits to start: five dollars a transaction, twenty a day, fifty a month. Not because that's the budget — because that's the discipline. You raise them only when the gate has earned it.

Here's why this matters, and why the order matters: treasurer before wallet.

The industry wants agents that spend. It hasn't built agents that deserve to.

The agent-commerce conversation has converged on a strange consensus: the milestone is the transaction. Demos end with the applause line — "and it booked the flight," "and it paid the invoice" — as if the ability to move money were proof of intelligence. It isn't. Moving money is proof of access. Access without judgment is not autonomy. It's a loaded credential.

The opposing camp isn't better: human approval on every cent. A tap-to-confirm for a four-dollar API call. That doesn't make the agent safe; it makes it a puppet with expensive latency, and it teaches the builder nothing about what the agent would actually do unsupervised. The moment you look away, the puppet has no reflexes of its own.

Both sides skip the real engineering question: what does the agent do with money when nobody is watching?

At N=1, the wallet is mine

I'm building for one user: me. That fact, which has organized this whole series — the right to refuse, the Tuesday audit, the deletion discipline — reaches its sharpest point here. There is no crowd to hide a bad charge in. Every mistaken dollar comes out of my pocket, personally, the same afternoon. No averaging, no "acceptable loss rate," no risk department. Just me, checking a statement.

That concentrates the mind wonderfully. It also clarifies what a spending gate actually is. It's not bureaucracy and it's not distrust — it's the same principle as everything else in this stack: the agent's freedom is defined by what it can refuse, audit, and confess. Money is just the domain where those abstractions become concrete. A refusal mechanism that can't say no to a payment is a demo feature. An audit that doesn't cover spend is theater. Honest degradation that doesn't apply to financial calls is a lie by omission.

So the treasurer completes the set. Every capability in this system already has to survive three interrogations — can it refuse, does it survive a random Tuesday, will it confess when it's degraded. Now every paid call gets a fourth: can it survive the treasurer.

The mechanics, because principles without mechanics are slogans

The gate runs before everything else — head of the guardrail chain, not an afterthought bolted on at the end. Its rules are deliberately dumb, because dumb rules are auditable:

  1. Whitelist first. Unknown payee, no payment. Not "review" — no. New relationships get added deliberately, by a human, on purpose.
  2. Three caps, not one. Per-transaction, per-day, per-month. A single cap is a speed bump; three caps are a shape. They force the question "is this pattern normal?" instead of just "is this charge small?"
  3. Payee-change detection. Same label, different destination account than history shows? That's the oldest trick in fraud, and it gets the strictest treatment: forced escalation to human review, every time. No learning it away, no "smart" exceptions.
  4. Fail closed. Unconfigured, uncertain, or erroring — the gate denies. An open-by-default treasurer is a decorative treasurer.

And the escalation path matters as much as the denials. "Review" is not a synonym for "eventually yes." It routes to a human queue with the full context — what was attempted, why it tripped, what the history looks like — and the final word on money stays human. That's not a limitation of the design. That's the design.

The uncomfortable corollary

Here's what building this taught me: you don't actually believe your agent is autonomous until you've given it the power to cost you money and then watched what it does. Most "autonomous agent" demos are choreography — the wallet is empty, the card is fake, the stakes are zero. Put a real five-dollar cap in front of it and suddenly every design decision gets honest. The refusal logic gets tested. The audit trail gets read. The degradation labels get checked, because now a degraded answer can degrade into a degraded charge.

So here's my rule, running in production this week: no wallet without a treasurer. No spending capability ships until the gate that governs it is already live, tested, and default-deny. Autonomy is not the absence of constraints — it's constraints you built on purpose, before you needed them.

Stop giving your agent a wallet. Give it a treasurer first. The demos will be less exciting. The statements will be cleaner. And the first time the gate says no to something you would have approved while distracted, you'll understand what the module was really for: it wasn't protecting the money. It was protecting the builder from his own inattention.

Me. The builder is me. The wallet is mine. The treasurer works for me, not for the agent — and that's exactly the point.

Top comments (0)