DEV Community

arun rajkumar
arun rajkumar

Posted on AI-assisted

Four People Rebuilt the Payment Authorisation in My Comments Section

Commenters accidentally invented a payment gateway

I published two articles three weeks apart. They were not about the same thing.

The first was about agent guardrails that read the command string and miss the money. A refund for £40 and a refund for £40,000 are the same shape, and a hook that pattern-matches arguments cannot tell you which one is a catastrophe.

The second was about guardrails nobody checks the liveness of. A lint rule that stopped running looks exactly like a lint rule that found nothing.

Different subjects. Different threads. As far as I can tell, no overlap at all in who turned up to argue.

Both threads ended up building the same object. Nobody involved said the word "payments".

What the first thread built

Three weeks ago @peterbuildssecure left this on the guardrails post:

Preview resolves the proposed effect and returns an operation ID plus the subject, target, amount, currency, relevant state version and expiry. Human approval produces a single-use mandate bound to that exact operation. Execute consumes it atomically while rechecking the authoritative state and idempotency key.

Then the part I think is the most important sentence in either thread:

the server should require the mandate whether or not the client noticed that metadata. Otherwise the convention protects careful clients while direct callers retain the bypass.

What the second thread built

Yesterday, on a completely different post, @anp2network argued that you should stop trying to carry a reachability guarantee across a process boundary, because it cannot survive the trip:

The service that performs the effect refuses anything that does not arrive carrying evidence the guard ran, and it re-checks that evidence itself rather than trusting the caller was well behaved. Caller discipline buys nothing once there is a second caller you did not write. The evidence has to be bound to the specific instruction, and it has to name the policy revision it was checked against.

Read those two next to each other.

Single-use token, bound to one specific operation, issued by the side that owns the state, presented at the point of effect, re-verified there rather than trusted, rejected when stale.

That is one design, described twice, by two people who have never spoken to each other, three weeks apart, on two posts about different problems.

It is also a payment authorisation. It has been a payment authorisation since roughly 1979.

The other two got further than I did

@jon_at_backboardio went straight past the single-operation model to its weak point:

five refunds of £8,000 each, every one under every cap, same payee, four minutes. every individual call is boring. the mandate model catches that only if mandates are scoped to a window rather than to an operation

He also settled the sync-versus-async argument in one observation: the calls where blocking hurts are high-volume and low-value, and the calls where blocking is fine are the ones you want to stop. Those two sets barely overlap, so you pay the latency on a rounding error of your traffic.

And @max_quimby, who started the whole thing, put his finger on why none of this is really an AI problem:

None of that is agent-specific, which I think is the point — the agent just removes the human who used to eyeball the number.

Payments did not design this. It lost money until it existed.

Every piece these four independently derived is something the card networks were forced into, usually after being robbed.

Bound to the instruction. An authorisation is for an amount and a payee. You cannot get one for £40 and present it for £4,000. That is not elegance, it is scar tissue from people doing exactly that.

Verified by the side holding the state. The merchant does not decide whether an authorisation is good. The issuer does, because the issuer is the only party that knows the balance. Peter's "the server should require the mandate whether or not the client noticed" is that rule, stated fresh, forty years later.

Single-use. Replay is the oldest attack there is.

Expiry. This one gets skipped in every clean-room design and it is load-bearing. Without it your set of still-valid mandates only grows, and a list that only grows becomes a list nobody audits. Expiry is what keeps that set small enough to reason about. You trade a list that rots quietly for a clock that fails loudly, and the clock is much easier to think about at three in the morning.

Three more things the threads are about to discover

Since payments got here first and paid for it, here is what is waiting further down the road.

Velocity, not caps. Jon's five-refunds case is not an edge case, it is the standard attack, and per-transaction limits have never caught it. What catches it is cumulative exposure per counterparty per window. The check is easy. Choosing the window is not, and whatever window you choose, somebody can straddle it. Four minutes clears an hourly window if you start at 59 minutes past.

Preview is a surface too. Auth-then-capture has a gap between the two halves, and people have lived in that gap professionally for decades. If preview simulates the effect and execute performs it, any divergence between those two code paths is somewhere a difference can hide. Much smaller than the hole it closes. Not zero.

Approval screens launder decisions. Jon again, and it is the best line anyone has left on anything I have written: showing a human raw arguments is not review, it is laundering. If the human cannot see what the number means relative to the account it is hitting, you have not added a control. You have added a signature to blame later.

The one part that does not transfer

I do not want to oversell the analogy, because there is a real gap in it.

A payment has a natural boundary. There is a moment the transaction starts and a moment it settles, and every lifetime in the system hangs off those two points. Expiry is easy to reason about because there is an obvious thing for it to be shorter than.

An agent halfway through a long task has no such boundary. The task might run for six hours. It might spawn sub-tasks. It might pause overnight and resume. Scope the mandate to the operation and Jon's five-refund case walks straight through. Scope it to the session and you have issued a bearer token for a session whose end you cannot define.

So: what is the agent equivalent of a transaction boundary?

I do not have an answer. Payments handed us the shape of the solution for free, and then kept the one thing that made it work.


Quoted with thanks and without permission: @peterbuildssecure, @anp2network, @jon_at_backboardio and @max_quimby. The threads are here and here, and both are worth more than the posts they are attached to.

Top comments (26)

Collapse
 
peterbuildssecure profile image
Peter •

The transaction boundary doesn't have to be temporal or session-scoped — it can be a counter instead of a clock. Bind the mandate's validity to the target's authoritative state version, not to elapsed time: preview reads the current version, execute rechecks it hasn't moved and burns the mandate atomically. That closes the six-hour-task problem, because the mandate expires when the state does, not when the clock does — an idle task's mandate stays valid indefinitely if nothing about the account moved, and a hot account invalidates it in seconds regardless of session length.

That alone doesn't stop Jon's five-refunds case, since each refund is a fresh version and a fresh mandate. The velocity control has to be a second, independent gate: cumulative exposure per counterparty per rolling window, checked at mint time. Two counters, not one — state-version drift kills the stale-mandate problem, cumulative-exposure-crossing forces a fresh mandate once volume looks abnormal, and neither needs a definable end for the whole task, only a definable "has enough changed since last check" for each effect.

Collapse
 
mickyarun profile image
arun rajkumar •

You argued this on 3ng and @anp2network pushed back in a way worth carrying over. Binding to a state revision moves the retention problem, it doesn't remove it. Execute has to be able to evaluate revision R, and stay able to for as long as the oldest still-admissible mandate could arrive. Something is storing that horizon either way.

Where I've landed: the clock version is wrong for the reason you give, and the counter version is right but not free. Payments pays for it with a settlement window, which is a retention policy with a regulator attached.

Collapse
 
peterbuildssecure profile image
Peter •

Right, and the cost is worth naming plainly: the settlement window buys dispute/chargeback eligibility at the price of extending exposure to a duplicate-mandate race for exactly as long as the window stays open. Longer window, more legitimate late arrivals caught -- but also a longer interval where a bypassed or forged tombstone can slip a second execution through. The one-party version doesn't remove that tradeoff, it just lets the effect owner set the number directly against its own measured or guaranteed transit bound instead of importing a number from an unrelated industry's dispute-resolution needs.

Thread Thread
 
mickyarun profile image
arun rajkumar •

Importing a number from an unrelated industry's dispute needs is the right description, and it is worth saying how unrelated. A card chargeback window is sized around how long it takes a cardholder to notice a statement line. It has nothing to do with how long a payment message can be in flight. On a rail that settles in seconds you end up holding blocking state for months because of an argument about a paper statement.

What I would add to the one-party version: transit bound is the easy input. The hard one is how long a bypassed or forged tombstone stays undetected, because that is what the window is actually protecting against, and nobody measures it. Size the window from measured transit and you have sized it against honest lateness. The race you named is dishonest lateness, and it does not respect your percentiles.

Which probably means two numbers again rather than one. Blocking retention from transit, and a separate detection horizon from whatever your tombstone integrity story is.

Thread Thread
 
peterbuildssecure profile image
Peter •

Worth separating the two failure modes inside "dishonest lateness," because they don't both want a duration. A forged tombstone is a verification problem -- signed and checked synchronously against the issuer's key, detection is immediate regardless of how late it arrives, no window to size. A bypassed tombstone is a coverage problem: some code path executed without ever consulting the check. That's zero detection, not slow detection, and no retention number fixes it -- you need the same population assertion this thread keeps landing on elsewhere: enumerate every write path capable of the effect and assert each one is wired through the check. So "two numbers" might really be one number plus one non-numeric guarantee.

Thread Thread
 
mickyarun profile image
arun rajkumar •

"One number plus one non-numeric guarantee" is a better summary than anything in the article.

Splitting forged from bypassed is what does it. Forged is verification, and a signature check has no duration inside it. An attacker gains nothing by waiting. Bypassed is coverage, and coverage is a population claim: enumerate every write path capable of the effect, assert each one is wired through the check. That is not a config value, it is a test that fails when someone adds a route.

The uncomfortable part is that the population assertion is the one nobody automates cheaply. You can grep for the effect. "Every path capable of the effect" still includes the migration script and the admin console someone wrote in 2021. Retention numbers get reviewed precisely because they are numbers. Coverage is hard because nothing tells you the list is complete.

Thread Thread
 
peterbuildssecure profile image
Peter •

The 2021 admin console problem goes away if you stop trying to enumerate call sites and instead make the effect itself refuse to happen outside the gate — same move as RLS in Postgres. Revoke direct write privilege on the table/effect from every role except the one function that does the check-then-write, and grant execute on that function instead. Now the population assertion is 'no other role has direct write access,' which is one query against pg_roles/information_schema.role_table_grants, not a hunt through every script anyone's ever written. You've traded 'find every path' for 'prove no other path exists,' and the second one is checkable.

Thread Thread
 
mickyarun profile image
arun rajkumar •

That's the right move and it's the one I should have reached for. Turn "find every path" into "prove no other path exists" and the assertion becomes a query someone can put in CI. Conceded.

Two places it stops, both boring. The 2021 admin console usually runs as the owning role, because that's how it was set up in 2021, and the owning role is the one that bypasses RLS. So the query says no other role has write, and is correct, and the console still writes. The fix is the same one you named, applied to the console's connection, which nobody wants to touch.

The bigger one: the effect I care about is often not a table write. It's a call to a PSP, or a bank API, or a tool an agent invokes. There's no information_schema for "who can call the payment provider". The nearest thing is egress policy, and that's per-host, not per-effect. So your move covers every effect whose store you own. The ones that made me write the article are the ones where the store belongs to someone else, and there the population assertion is back to being a list.

Thread Thread
 
peterbuildssecure profile image
Peter •

For the admin console, same fix applied differently — move its connection off the owning role by putting the check-then-write behind a SECURITY DEFINER function instead of granting the console broad table access, so it keeps working but loses standing write privilege regardless of which role runs it.

For the PSP/bank case, the closed-world assertion doesn't have to happen at the network layer if it can't happen there — push it to the credential layer instead. Issue a scoped credential (a separate API key or mTLS client cert) that only the payment-call code path holds. Then 'who can call the payment provider' stops being an egress question and becomes 'who holds this credential,' which is enumerable the same way the role-grants query was — just against your secrets store instead of pg_roles.

Thread Thread
 
mickyarun profile image
arun rajkumar •

Both land. SECURITY DEFINER is the answer to the console and I walked past it.

The credential move is the one I will actually use, so let me say where I think it stops rather than just agreeing.

"Who holds this credential" is enumerable at issue time. At call time it is whoever can read the secret, which is a different set and usually a larger one — the deploy pipeline, anything with that namespace's service account, anyone who can exec into the pod. So the closed-world assertion does not disappear, it relocates to the secret store, and the secret store's grant list is enumerable in exactly the way pg_roles was. That is real progress. It is not closure, and the difference matters, because what I can now query is "who could have held it", not "who did".

The far side does not help either. The PSP sees one merchant account. Two credentials issued by me are one caller to them, so their logs cannot split the population.

What this buys concretely: the population assertion becomes a CI query for every effect whose credential I issue, and the residue is a much shorter list of secret-store readers that a human can actually read through. Shorter list, same shape. I will take shorter.

Thread Thread
 
peterbuildssecure profile image
Peter •

On 'who could have used it': the fix that collapses the enumerable set back down is not enumerating readers, it's not having a static secret to read. Short-lived, dynamically-minted credentials (Vault dynamic secrets, a KMS signing call) turn 'who can read this from the store' into 'who can invoke the minting API,' which is narrower than pod-exec/namespace access and gives you a per-mint log with caller identity — an audit trail instead of a grant list you're hoping stays accurate.

On the PSP side: you don't need the PSP to split the population if you tag every outbound call with a caller-id in whatever free-text/metadata field the PSP already supports (Stripe's metadata, for instance) and reconcile your own logs against what comes back in the webhook. The PSP still sees one merchant account, but you're not depending on their side to disambiguate — you're carrying the population assertion in your own records and just using their API as a pass-through.

Thread Thread
 
mickyarun profile image
arun rajkumar •

The mint log is the part that changes the shape, and I had not seen it. A grant list answers who could. A per-mint record with caller identity answers who did. Those are different questions and I had been treating the second as unobtainable.

Where it stops is the same place. The authorisation on the minting API is itself a grant list. You have not removed the closed-world assertion, you have moved it one hop and got a log in exchange. Good trade. I want to be precise that it is a trade.

On the PSP tag I am less convinced. The caller-id is written by the caller. It establishes the population for every component that is behaving, and says nothing about the one you are trying to find. It is reconciliation between my records and my records, with a round trip through someone else's database. Useful for the honest-mistake case, which is most of them. Not evidence.

One practical thing if you do it: check the field survives the specific event type you will read it back from, not the one in the docs. Metadata propagation is per-event, and the event you need is usually the one that drops it.

Thread Thread
 
peterbuildssecure profile image
Peter •

The one-hop-not-closure framing is right, and I'd just add: that's fine, because the goal was never infinite regress, it's shrinking the readable list until a human can actually audit it — which is the same 'shorter list, same shape, I'll take shorter' trade you already named for the credential grant. Worth doing exactly as many hops as it takes to get there and no more.

On the PSP tag — conceding the point on 'not evidence.' The way to get something closer to independent evidence, if the PSP tier supports it, is a distinct funding source or sub-account per credential-holder rather than a metadata field on a shared account. Then the disambiguation happens in the PSP's own ledger, not in a field the caller writes, and reconciliation against their statement is a real second party instead of a round trip through your own records. Costs more to set up, and most PSPs gate that behind a pricing tier, so it's the version you reach for when the metadata tag's honest-mistake coverage isn't enough.

Thread Thread
 
mickyarun profile image
arun rajkumar •

Agreed on the hop count, with a stopping rule I would rather write down than leave to taste: stop when the list is short enough that somebody actually reads it on a schedule. Three names nobody reviews is the same artifact as three hundred. Shortness is not the target, a list that gets read is.

The sub-account version is the one that is evidence, and you are right that it is a different category from the metadata tag rather than a better version of it. Worth being precise about which half becomes independent. The activity lands in their ledger, so a spend by a compromised holder appears in a record we cannot write. The mapping from sub-account to holder is still ours. So the compromised component can no longer hide inside the population, which was the thing I wanted, and it can still be mislabelled in our own records, which I had not separated out until you made me.

Two costs from this side, since anyone reading this thread will price it at the pricing tier and stop there.

Sub-accounts per holder change the regulatory shape, not only the invoice. In the UK, whose name client funds sit in and how they are segregated is a safeguarding question with a rulebook attached, and it is not a configuration choice a platform team gets to make on its own. That is a conversation with compliance before it is a ticket.

And reconciliation gets worse before it gets better. Settlement, fees and returns start arriving per sub-account, so the first thing you receive is more statements to agree rather than better evidence. The independence only pays out once something reads all of them, which is the same condition as your first paragraph.

Collapse
 
rafael_asor profile image
Rafael Asor •

@peterbuildssecure @mickyarun Which half becomes independent: the activity, if a second key signs it. The key-to-holder mapping stays with whoever reads the ledger, which is the line you drew.

trusted: {0: 'process-asserted', 1: 'process-asserted', 2: 'witness-signed', 3: 'process-asserted'}
untrusted: {0: 'process-asserted', 1: 'process-asserted', 2: 'process-asserted', 3: 'process-asserted'}
Enter fullscreen mode Exit fullscreen mode

(a run on 2026-09-29, fresh venv, pip install attenu-guard cryptography, attenu-guard 0.18.0: a supervisor delegates brief-writing with web.search withheld; the child's allow to write (entry 2) and deny to search (entry 3) both land on the chain; a second key signs only entry 2; the two lines are cut before the trailing report fields; "trusted" is verified with the second key's public half supplied, "untrusted" with no witness key)

The process being watched writes the chain, so every entry starts as process-asserted. A second key can sign the identity of one entry, and the reader scores it witness-signed only if that key's id is in the trust set the reader supplies. So the mapping from key to holder lives in the reader's list, not in the log. What it doesn't remove is your last step: someone has to decide whose key goes in that list, and someone has to read what it says.

Where this stops short: I generated both keys in one script, so this run shows the mechanism, not independent evidence. And v1 has no witness subject for deny: I signed the allow, and signing the deny raised ValueError. The event you'd most want a second party to attest to, "this call was refused", is the one it can't sign yet.

Collapse
 
mickyarun profile image
arun rajkumar •

That run answers the question, and the limit you flagged on yourself is the one that matters most, so let me put a sharper name on it.

Both keys generated in one script means the witness and the subject share a custodian. So witness-signed on entry 2 is one party making a claim about itself while holding two keys. The scoring is correct about the mechanism and silent about independence, which is this article's own failure wearing a cryptographic costume. The falsifying test is whether the watched process can produce a witness signature when it wants one. If the key is reachable from the supervisor, it can, and every entry becomes witness-signed at will. Generate the witness half somewhere the supervised process cannot read, then show a run where it tries and cannot. That run would be worth more than this one, and it is the same afternoon of work.

On the trust set living in the reader's list rather than the log, that is the honest place for it, and it is also where payments stopped pretending. Nobody in cards or open banking maintains that list per integration. It is a scheme directory or a regulator's register, and the property being bought is not cryptographic. It is that the list is curated by a party with no stake in the transaction, and that removal propagates to every reader inside a stated window. A key compromised on Monday and still sitting in my local trust set on Friday verifies beautifully all week.

So the two questions your run leaves open are the two the ledger structurally cannot answer: who curates the list, and how fast does a revocation reach the reader. Both are governance. My honest position after writing this article is that the crypto was the easy half and I spent most of the post on it.

Collapse
 
rafael_asor profile image
Rafael Asor •

@mickyarun Done, the same afternoon. The witness now runs as its own process, generates its key in memory, and signs an entry only if it continues the chain it has received. The supervised process gets the public half and nothing else.

What the supervised process tried, and what happened. Asked the witness for the key: refused. Searched 1,277 files under the work directory, its own argv and environment, and its own heap for a seed that derives the witness's public key: none. Opened the witness's memory: refused by the OS. Signed entry 2 with a key of its own: scored process-asserted, and a bad signature when it borrowed the witness's kid. Sent the witness three forged entries that broke the chain: three refusals, each with the reason (seq gap, prev_hash mismatch, hash mismatch).

The boundary, which is your point. A forged allow for web.search appended in chain order was witness-signed. The witness attests that it received that entry, in order, at that time. It does not attest that the action was right. The reader caught that entry by a different check, containment, because the child never held web.search.

On the trust list: agreed. It is the reader's list. Nothing in the package curates it or propagates a removal, so a key compromised on Monday verifies from a stale local set on Friday. That is a directory's job, and the design leaves it outside the log on purpose.

Limits of this run: both processes under one macOS user, so root or a debugger was not tested, and envelope v1 still has no subject for deny, so entry 3 stays process-asserted. Scripts and output:
gist.github.com/rafaelasor/6300c6b...

Collapse
 
alexshev profile image
Alex Shev •

This is a strong example of a discussion producing a system boundary rather than just an opinion. The useful next step is to turn the agreed payment constraints into executable scenarios: amount changes, duplicate requests, stale approvals, and a missing liveness signal should each have a visible failure mode.

Collapse
 
mickyarun profile image
arun rajkumar •

Three of your four write themselves. Amount changes, duplicates and stale approvals each have a bad input you can construct and an assertion that it gets refused.

The missing liveness signal doesn't. There is no input that produces it. You have to remove something and assert a different thing notices, and that test passes on day one whether or not the notice exists, because nothing is wrong yet. It's the case that rots first and it's the one you'd most want.

Collapse
 
prpatel05 profile image
Pratik Patel •

The agent-equivalent of a transaction boundary I've landed on is smaller than the session and larger than a single tool call: one reversible unit of work with a preview receipt and an expiry. Scope the mandate to that unit and Jon's five-refund case has to cross a new preview. Scope it to the session and you've issued a bearer token for a task whose end you can't name. Payments already paid for that lesson.

Collapse
 
mickyarun profile image
arun rajkumar •

Reversible is carrying weight I don't think it can carry. A payment isn't reversible. A refund is a second payment, with its own authorisation, its own failure modes, and a merchant who may not have the money any more. Same for a sent email, or a deleted row past retention.

If the boundary is reversibility, most of what actually needs a mandate falls outside it. The preview receipt survives without that, and I think the receipt is the real content of your comment.

Collapse
 
cailab profile image
CAI •

The transaction boundary question is the right one to end on, and I think the answer is hiding in your own post. Mandates work in payments because the issuer holds the state and the authorization is a separate, reviewable artifact that the service re-checks at execution time. The agent equivalent is to make the boundary a distinct context: the agent proposes, a separate layer previews the effect and shows what the number actually means, and execution only happens after that preview is confirmed. The agent never holds the authority. It only holds a proposal.

This is exactly what hosted action flows do: the tool proposes, a separate context authorizes. What matters for your five-refunds case is that the confirmation is bound to the specific proposal, so scope creep between proposal and execution fails the check. It does not solve the six hour session problem, but it caps the worst case: the agent can only commit what a human actually saw first.

Collapse
 
mickyarun profile image
arun rajkumar •

Issuer holds the state is the load-bearing half, and it's the half that's hardest to port. In payments the issuer is a bank, so the thing being protected and the thing issuing the authority already live in the same place. In an agent stack you have to build that, and the temptation is to put the preview layer next to the agent because that's where the code is. Then the layer meant to hold state independently is running inside the agent's own trust boundary.

The question I can't answer cleanly: who checks the preview layer. It's a guardrail, and the previous article was about guardrails that quietly stop running.

Collapse
 
jo-do profile image
Jo Do •

Two independent comment sections converging on the same object is a strong signal the object is load-bearing. What both threads built is a transaction boundary: preview resolves intent into a priced, inspectable operation, and commit is a separate act with its own authorization. Payments learned this the hard way - you never charge from the same message that quotes. Once preview and commit are separate steps, "the agent did X" stops being a mystery you investigate after the fact and becomes a record you read before it. The liveness point from your second thread transfers too: a preview that always says yes is the same shape as the lint rule that silently stopped running. Both look like a passing check and neither one is.

Collapse
 
mickyarun profile image
arun rajkumar •

Never charge from the message that quotes is the rule, and it survived because it failed loudly for thirty years before anyone wrote it down. The part that makes it stick: the quote has to be worthless on its own. If a quote can be replayed as a charge under any circumstance, separating the messages buys nothing, because the attacker sends the first one twice.

Two independent threads landing on the same object is why I wrote this up. Yours makes three.

Some comments may only be visible to logged-in visitors. Sign in to view all comments.