If you're building anything that lets an AI agent act on a user's behalf — send a message, book something, control a device — there's a failure mode that's easy to miss until it bites you: the gap between "permission was granted" and "permission is still valid right now."
The scenario
Imagine this timeline:
T+0:00 — An agent has permission to perform an action, and that authority is valid.
T+0:04 — The authority is revoked, or the world changes underneath it (a contact's number updates, a calendar slot fills, a setting reverts).
T+0:05 — The agent still has the technical capability to execute.
T+0:06 — A downstream system records the action as complete.
The agent had permission. Was it still authorized at the moment it actually mattered? Those are different questions, and most systems don't distinguish them.
Why "blocked" often lies
Say you catch the problem and try to stop the action after authority changes. If the request has already been dispatched to an external system — an SMS provider, a calendar API — simply flipping your own app's state to "blocked" doesn't prove anything happened, or didn't. You're guessing, dressed up as a system that checked.
This showed up for me building a confirm-before-execute layer for phone commands. A stale result from an earlier attempt rendered as "executed" right after a new attempt correctly denied the same action. The denial was truthful about the decision. It was silent about the actual state of the world.
Three states, not two
Most systems track two outcomes: executed, or blocked. That's not enough. The honest model needs three:
EXECUTED — observed evidence the intended side effect actually occurred.
DENIED_CONFIRMED — denied, and verified the action never crossed the execution boundary.
DENIED_UNRESOLVED — denial was intended, but available evidence can't independently establish whether a downstream side effect occurred anyway.
That third state is the one most systems don't have a slot for at all. Without it, uncertainty silently gets rounded up to a confident "blocked" — which is worse than an honest "I don't know," because it actively points you away from checking.
The fix that keeps showing up
Talking this through with people building completely different systems — WordPress admin tools, AI ops platforms, agent authorization layers — the same pattern kept surfacing independently:
Never trust an earlier authorization check. Re-verify right before execution, not when the action was first requested.
Bind confirmation to exact parameters. A dry run returns a short-lived token tied to the specific values (recipient, time, amount). Confirm has to present that token, byte-identical, or it's treated as a new plan requiring fresh approval — not a rubber-stamp of the old one.
Once dispatched externally, "blocked" isn't a real status anymore. Only a provider receipt or webhook can close the loop. Until then, the honest label is "submitted, unresolved" — not "blocked," and definitely not "executed."
What this changes practically
For a phone-command interface, the gap between showing someone a plan and firing it is usually seconds — but "short" isn't "zero." A dry run now returns a hash of everything the action depends on. Confirm has to match that hash exactly, or it's rejected as stale and requires a fresh confirmation. Cheap to build, and it closes a real gap that used to be invisible.
The broader lesson: if your system can't tell "we decided not to do this" apart from "we don't know what happened," you don't have a permission system — you have a permission system's confident-sounding guess.
Building StareBrain — natural language commands for Android, with exactly this confirmation model at its core. Currently pre-launch: starebrain.vercel.app
Top comments (5)
The third state is the right call, and the argument for it — that rounding uncertainty up to "blocked" actively points you away from checking — is the part worth quoting at people. Two pushes.
DENIED_UNRESOLVEDonly helps if something resolves it. If it's a terminal value in your log then you've renamed the uncertainty rather than removed it, and within a few weeks it becomes the status everyone's eyes slide past. It needs a reconciliation pass that owns it: a job that re-queries the provider for unresolved records and promotes each one to EXECUTED or DENIED_CONFIRMED, plus an age at which failing to resolve escalates rather than quietly ages out. Otherwise the honest label decays into exactly the noise that the confident "blocked" was.Your two mechanisms cover different failure classes, and your own example shows the gap. Binding the confirm token to exact parameters catches parameter drift: the recipient changed between plan and confirm. But in your T+0:04 list, "a calendar slot fills" is world drift — the parameters are byte-identical, the token validates cleanly, and the plan is stale anyway. So token binding doesn't subsume re-verification and re-verification doesn't subsume token binding. A system with only the token feels rigorous and still books the taken slot.
The generalization I keep running into is one layer further down, on-device.
Result.success()from a WorkManager job, or a 200 from a local write, is your own state rather than observed evidence that the side effect happened. Same three states apply and the "did it cross the boundary" question is identical; it's just that the boundary is a filesystem or a content provider instead of an SMS gateway. Any boundary you can't cheaply read back across needs the third state.All three of these land, and the second one is the one that actually breaks something I thought was solved.
On reconciliation: right, DENIED_UNRESOLVED as a terminal label is just a renamed uncertainty, not a resolved one. Needs an active job re-querying the provider on a schedule, promoting to EXECUTED or DENIED_CONFIRMED when evidence arrives, and an explicit age-out that escalates (surfaces to the user, or to me, as a real problem) rather than silently expiring into "probably fine." Don't have that built yet — this is the concrete next piece.
On token binding vs. re-verification: this is the sharper catch. I'd been treating the parameter-bound token as covering "did the world change," when it only covers "did the plan change." Your calendar example is exact — identical parameters, valid token, stale plan, and the token has no way to know. These are genuinely separate mechanisms solving separate failure classes, and I was one sentence away from believing the token alone was sufficient. It isn't. Re-verification against live state has to run independently, right before dispatch, regardless of whether the token validates.
On the on-device generalization: hadn't extended the model that far, and it's obviously the same shape once you say it. Result.success() from WorkManager is exactly as much "my own claim about myself" as a provider's webhook is — I'd been mentally reserving DENIED_UNRESOLVED for external boundaries and treating local writes as trustworthy by default, which is the same unearned confidence this whole piece is about, just one layer closer to home. Any boundary I can't cheaply read back across needs the third state, full stop, whether it's a network call or a local content provider.
Going to go audit every boundary in the app against that test, not just the ones that felt risky by intuition.
The age-out escalation is the piece I would build ahead of the re-query job. It is the smaller of the two, and it is the one that stops the failure mode being silent.
One structural suggestion on its shape: have the record carry a deadline set at creation, rather than the reconciler evaluating a timeout at read time. A deadline derived on read quietly slides forward every time nobody looks, so records age indefinitely without ever aging out — which looks identical to a healthy queue. Past the deadline the record moves to a state a human sees, and deliberately not a terminal one:
DENIED_UNRESOLVEDearning a terminal label was the original bug, and an age-out that expires into a different terminal label just reintroduces it one layer down.The other thing worth wiring early is making the re-query job's own health visible separately from the records it reconciles. A reconciler that has been failing for a week and a week with genuinely nothing unresolved render the same dashboard, and the natural reading is the flattering one. Emit "reconciler ran, examined N, promoted M" — including when N is zero — so the absence of alerts is evidence rather than an assumption.
On token binding versus re-verification: does your provider expose an idempotency key you can replay? If a replay returns the original result rather than executing again, part of "did the world change" collapses into a cheap read, and full re-verification is only needed for the parameters the provider does not bind. That will not cover the stale-calendar case, which is genuinely the other mechanism — but it usefully shrinks how often you have to reach for it.
Deadline-at-creation over timeout-at-read is the right fix, and it's an embarrassingly common trap I'd have walked straight into — a computed timeout that silently resets every time nobody checks is functionally identical to no timeout at all, just with better optics. Setting the deadline once, at write time, and having age-out move the record to a visible, non-terminal state (not a new "expired" terminal label wearing a trenchcoat) is the right shape. Noted, building it that way.
The reconciler-health point is the one I hadn't even considered a separate problem. "Zero unresolved" and "reconciler dead for a week" render identically on a naive dashboard, and I'd have built exactly that naive dashboard. Emitting a heartbeat — ran, examined N, promoted M, including N=0 — turns silence into a signal I can actually check, instead of an assumption I'd have been making by default.
On idempotency keys: yes, worth checking per-provider, and you're right that it only shrinks the surface rather than replacing re-verification. Where a replay-safe idempotency key exists, a retry or duplicate dispatch collapses into a cheap read instead of a live re-check — genuinely useful for the "did I already do this" failure mode. But you're right it's orthogonal to the calendar case: an idempotency key tells me "this exact request already executed," not "is this request still valid to execute," which is precisely the plan-vs-world distinction from your first comment. Two separate cheap wins, neither one covering the other.
Some comments may only be visible to logged-in visitors. Sign in to view all comments.