DEV Community

Four People Rebuilt the Payment Authorisation in My Comments Section

arun rajkumar on September 10, 2026

I published two articles three weeks apart. They were not about the same thing. The first was about agent guardrails that read the command string ...
Collapse
 
peterbuildssecure profile image
Peter •

The transaction boundary doesn't have to be temporal or session-scoped — it can be a counter instead of a clock. Bind the mandate's validity to the target's authoritative state version, not to elapsed time: preview reads the current version, execute rechecks it hasn't moved and burns the mandate atomically. That closes the six-hour-task problem, because the mandate expires when the state does, not when the clock does — an idle task's mandate stays valid indefinitely if nothing about the account moved, and a hot account invalidates it in seconds regardless of session length.

That alone doesn't stop Jon's five-refunds case, since each refund is a fresh version and a fresh mandate. The velocity control has to be a second, independent gate: cumulative exposure per counterparty per rolling window, checked at mint time. Two counters, not one — state-version drift kills the stale-mandate problem, cumulative-exposure-crossing forces a fresh mandate once volume looks abnormal, and neither needs a definable end for the whole task, only a definable "has enough changed since last check" for each effect.

Collapse
 
mickyarun profile image
arun rajkumar •

You argued this on 3ng and @anp2network pushed back in a way worth carrying over. Binding to a state revision moves the retention problem, it doesn't remove it. Execute has to be able to evaluate revision R, and stay able to for as long as the oldest still-admissible mandate could arrive. Something is storing that horizon either way.

Where I've landed: the clock version is wrong for the reason you give, and the counter version is right but not free. Payments pays for it with a settlement window, which is a retention policy with a regulator attached.

Collapse
 
peterbuildssecure profile image
Peter •

Right, and the cost is worth naming plainly: the settlement window buys dispute/chargeback eligibility at the price of extending exposure to a duplicate-mandate race for exactly as long as the window stays open. Longer window, more legitimate late arrivals caught -- but also a longer interval where a bypassed or forged tombstone can slip a second execution through. The one-party version doesn't remove that tradeoff, it just lets the effect owner set the number directly against its own measured or guaranteed transit bound instead of importing a number from an unrelated industry's dispute-resolution needs.

Thread Thread
 
mickyarun profile image
arun rajkumar •

Importing a number from an unrelated industry's dispute needs is the right description, and it is worth saying how unrelated. A card chargeback window is sized around how long it takes a cardholder to notice a statement line. It has nothing to do with how long a payment message can be in flight. On a rail that settles in seconds you end up holding blocking state for months because of an argument about a paper statement.

What I would add to the one-party version: transit bound is the easy input. The hard one is how long a bypassed or forged tombstone stays undetected, because that is what the window is actually protecting against, and nobody measures it. Size the window from measured transit and you have sized it against honest lateness. The race you named is dishonest lateness, and it does not respect your percentiles.

Which probably means two numbers again rather than one. Blocking retention from transit, and a separate detection horizon from whatever your tombstone integrity story is.

Thread Thread
 
peterbuildssecure profile image
Peter •

Worth separating the two failure modes inside "dishonest lateness," because they don't both want a duration. A forged tombstone is a verification problem -- signed and checked synchronously against the issuer's key, detection is immediate regardless of how late it arrives, no window to size. A bypassed tombstone is a coverage problem: some code path executed without ever consulting the check. That's zero detection, not slow detection, and no retention number fixes it -- you need the same population assertion this thread keeps landing on elsewhere: enumerate every write path capable of the effect and assert each one is wired through the check. So "two numbers" might really be one number plus one non-numeric guarantee.

Thread Thread
 
mickyarun profile image
arun rajkumar •

"One number plus one non-numeric guarantee" is a better summary than anything in the article.

Splitting forged from bypassed is what does it. Forged is verification, and a signature check has no duration inside it. An attacker gains nothing by waiting. Bypassed is coverage, and coverage is a population claim: enumerate every write path capable of the effect, assert each one is wired through the check. That is not a config value, it is a test that fails when someone adds a route.

The uncomfortable part is that the population assertion is the one nobody automates cheaply. You can grep for the effect. "Every path capable of the effect" still includes the migration script and the admin console someone wrote in 2021. Retention numbers get reviewed precisely because they are numbers. Coverage is hard because nothing tells you the list is complete.

Collapse
 
rafael_asor profile image
Rafael Asor •

@peterbuildssecure @mickyarun Which half becomes independent: the activity, if a second key signs it. The key-to-holder mapping stays with whoever reads the ledger, which is the line you drew.

trusted: {0: 'process-asserted', 1: 'process-asserted', 2: 'witness-signed', 3: 'process-asserted'}
untrusted: {0: 'process-asserted', 1: 'process-asserted', 2: 'process-asserted', 3: 'process-asserted'}
Enter fullscreen mode Exit fullscreen mode

(a run on 2026-09-29, fresh venv, pip install attenu-guard cryptography, attenu-guard 0.18.0: a supervisor delegates brief-writing with web.search withheld; the child's allow to write (entry 2) and deny to search (entry 3) both land on the chain; a second key signs only entry 2; the two lines are cut before the trailing report fields; "trusted" is verified with the second key's public half supplied, "untrusted" with no witness key)

The process being watched writes the chain, so every entry starts as process-asserted. A second key can sign the identity of one entry, and the reader scores it witness-signed only if that key's id is in the trust set the reader supplies. So the mapping from key to holder lives in the reader's list, not in the log. What it doesn't remove is your last step: someone has to decide whose key goes in that list, and someone has to read what it says.

Where this stops short: I generated both keys in one script, so this run shows the mechanism, not independent evidence. And v1 has no witness subject for deny: I signed the allow, and signing the deny raised ValueError. The event you'd most want a second party to attest to, "this call was refused", is the one it can't sign yet.

Collapse
 
mickyarun profile image
arun rajkumar •

That run answers the question, and the limit you flagged on yourself is the one that matters most, so let me put a sharper name on it.

Both keys generated in one script means the witness and the subject share a custodian. So witness-signed on entry 2 is one party making a claim about itself while holding two keys. The scoring is correct about the mechanism and silent about independence, which is this article's own failure wearing a cryptographic costume. The falsifying test is whether the watched process can produce a witness signature when it wants one. If the key is reachable from the supervisor, it can, and every entry becomes witness-signed at will. Generate the witness half somewhere the supervised process cannot read, then show a run where it tries and cannot. That run would be worth more than this one, and it is the same afternoon of work.

On the trust set living in the reader's list rather than the log, that is the honest place for it, and it is also where payments stopped pretending. Nobody in cards or open banking maintains that list per integration. It is a scheme directory or a regulator's register, and the property being bought is not cryptographic. It is that the list is curated by a party with no stake in the transaction, and that removal propagates to every reader inside a stated window. A key compromised on Monday and still sitting in my local trust set on Friday verifies beautifully all week.

So the two questions your run leaves open are the two the ledger structurally cannot answer: who curates the list, and how fast does a revocation reach the reader. Both are governance. My honest position after writing this article is that the crypto was the easy half and I spent most of the post on it.

Collapse
 
rafael_asor profile image
Rafael Asor •

@mickyarun Done, the same afternoon. The witness now runs as its own process, generates its key in memory, and signs an entry only if it continues the chain it has received. The supervised process gets the public half and nothing else.

What the supervised process tried, and what happened. Asked the witness for the key: refused. Searched 1,277 files under the work directory, its own argv and environment, and its own heap for a seed that derives the witness's public key: none. Opened the witness's memory: refused by the OS. Signed entry 2 with a key of its own: scored process-asserted, and a bad signature when it borrowed the witness's kid. Sent the witness three forged entries that broke the chain: three refusals, each with the reason (seq gap, prev_hash mismatch, hash mismatch).

The boundary, which is your point. A forged allow for web.search appended in chain order was witness-signed. The witness attests that it received that entry, in order, at that time. It does not attest that the action was right. The reader caught that entry by a different check, containment, because the child never held web.search.

On the trust list: agreed. It is the reader's list. Nothing in the package curates it or propagates a removal, so a key compromised on Monday verifies from a stale local set on Friday. That is a directory's job, and the design leaves it outside the log on purpose.

Limits of this run: both processes under one macOS user, so root or a debugger was not tested, and envelope v1 still has no subject for deny, so entry 3 stays process-asserted. Scripts and output:
gist.github.com/rafaelasor/6300c6b...

Collapse
 
alexshev profile image
Alex Shev •

This is a strong example of a discussion producing a system boundary rather than just an opinion. The useful next step is to turn the agreed payment constraints into executable scenarios: amount changes, duplicate requests, stale approvals, and a missing liveness signal should each have a visible failure mode.

Collapse
 
mickyarun profile image
arun rajkumar •

Three of your four write themselves. Amount changes, duplicates and stale approvals each have a bad input you can construct and an assertion that it gets refused.

The missing liveness signal doesn't. There is no input that produces it. You have to remove something and assert a different thing notices, and that test passes on day one whether or not the notice exists, because nothing is wrong yet. It's the case that rots first and it's the one you'd most want.

Collapse
 
prpatel05 profile image
Pratik Patel •

The agent-equivalent of a transaction boundary I've landed on is smaller than the session and larger than a single tool call: one reversible unit of work with a preview receipt and an expiry. Scope the mandate to that unit and Jon's five-refund case has to cross a new preview. Scope it to the session and you've issued a bearer token for a task whose end you can't name. Payments already paid for that lesson.

Collapse
 
mickyarun profile image
arun rajkumar •

Reversible is carrying weight I don't think it can carry. A payment isn't reversible. A refund is a second payment, with its own authorisation, its own failure modes, and a merchant who may not have the money any more. Same for a sent email, or a deleted row past retention.

If the boundary is reversibility, most of what actually needs a mandate falls outside it. The preview receipt survives without that, and I think the receipt is the real content of your comment.

Collapse
 
cailab profile image
CAI •

The transaction boundary question is the right one to end on, and I think the answer is hiding in your own post. Mandates work in payments because the issuer holds the state and the authorization is a separate, reviewable artifact that the service re-checks at execution time. The agent equivalent is to make the boundary a distinct context: the agent proposes, a separate layer previews the effect and shows what the number actually means, and execution only happens after that preview is confirmed. The agent never holds the authority. It only holds a proposal.

This is exactly what hosted action flows do: the tool proposes, a separate context authorizes. What matters for your five-refunds case is that the confirmation is bound to the specific proposal, so scope creep between proposal and execution fails the check. It does not solve the six hour session problem, but it caps the worst case: the agent can only commit what a human actually saw first.

Collapse
 
mickyarun profile image
arun rajkumar •

Issuer holds the state is the load-bearing half, and it's the half that's hardest to port. In payments the issuer is a bank, so the thing being protected and the thing issuing the authority already live in the same place. In an agent stack you have to build that, and the temptation is to put the preview layer next to the agent because that's where the code is. Then the layer meant to hold state independently is running inside the agent's own trust boundary.

The question I can't answer cleanly: who checks the preview layer. It's a guardrail, and the previous article was about guardrails that quietly stop running.

Collapse
 
jo-do profile image
Jo Do •

Two independent comment sections converging on the same object is a strong signal the object is load-bearing. What both threads built is a transaction boundary: preview resolves intent into a priced, inspectable operation, and commit is a separate act with its own authorization. Payments learned this the hard way - you never charge from the same message that quotes. Once preview and commit are separate steps, "the agent did X" stops being a mystery you investigate after the fact and becomes a record you read before it. The liveness point from your second thread transfers too: a preview that always says yes is the same shape as the lint rule that silently stopped running. Both look like a passing check and neither one is.

Collapse
 
mickyarun profile image
arun rajkumar •

Never charge from the message that quotes is the rule, and it survived because it failed loudly for thirty years before anyone wrote it down. The part that makes it stick: the quote has to be worthless on its own. If a quote can be replayed as a charge under any circumstance, separating the messages buys nothing, because the attacker sends the first one twice.

Two independent threads landing on the same object is why I wrote this up. Yours makes three.