Say a finance lead approves an AI agent to pay supplier invoices for the next 90 days. On day three, one of those suppliers gets flagged for fraud. The agent keeps running, and the approval record still says approved. Signed, in date, valid. Nothing in that record is wrong. It is out of date, and it has no way of knowing.
That is the problem I want to walk through this week, using our Agent Manifest.
An Agent Manifest is one signed record of everything that defines an agent when it is deployed: its prompt, its policies, its tools, its model, its data, its memory, its decision log, who it can hand work to, where it came from, and the humans who approved it. Anyone can check it later without trusting whoever runs the agent.
Human approval is one of those pieces. Our design decision on approvals checks four things: the person who approved is on the allowed list, the approval window has not run out, the approval is tied to this exact deployment, and the signature is real. The check runs offline. No call to an approval service, no dashboard to trust. Four yes-or-no answers from the record itself.
Where it stops working
All four can still be yes after the approval stopped making sense.
The record says a person approved something at a point in time. It does not say what they were relying on when they said yes, whether they can still stop the agent, or whether anybody objected. So in the invoice example the manifest verifies cleanly and the reason behind the approval is gone.
The expiry date does not fix this. An approval does not expire when the facts change on day three, and it does expire when nothing has changed on day ninety. Time is a stand-in for “still current”, and a weak one.
What we already shipped
Two small things, both on purpose.
First, PR 355, merged on 31 August, changed the spec to say plainly that an approval’s duration sets a window. It is not a promise that the approval is still current. No new field. It just stops people reading the field as something it never promised.
Second, PR 453, which shipped in agent-manifest 0.13.0 on 25 September. When all four checks pass, the verifier still says approved, and it now also says it cannot decide whether the approval still applies, because it has no evidence about the present. A result can be valid and undecided at the same time, and the README tells callers not to treat that as permission to go ahead. I like this one. The verifier now tells you what it does not know instead of staying quiet.
Who worked this out
This started as issue 348, opened by Ioana Valea on 27 August. The proposal was to record what the approver relied on, so a later check could notice when that changed.
Then @solloek369-arch raised the objection that changed the design. Everything in a manifest is fixed when it is issued; if a piece changes, you issue a new manifest. So comparing what the approver relied on against that same manifest will always match. It can never catch a change. What you compare against had to be decided first.
That split what an approver relies on into three kinds of thing: parts of the agent itself, facts about the outside world, and the approver’s own standing. Each needs a different check, and one status for all three would hide exactly the difference you care about.
Neither of them works for us. The person behind that objection described themselves as a snail on an F1 grid in this kind of work, and it is the reason the design changed. That is what an open problem is for.
Still open
Two things, and I would rather say so than wait for a tidy answer.
Deciding whether an approval still applies is open. The verifier can now say “I can’t tell”. Turning that into a real answer needs evidence about the present that the verifier does not have yet.
The words for “not established” are not settled, and that is bigger than this issue. trace-spec issue 279 is collecting the same question from six threads. A checker that cannot tell “this is false”, “this is unproven” and “nobody asked” apart will sooner or later be read as saying the first when it meant the third.
Try it on your own system
Take any approval your system records and ask three questions. What was the approver relying on, and would you find out if it changed? Who can stop this right now, and has anyone actually tried? Did anybody object, and where is that written down?
Most approval systems answer none of the three. None of them need cryptography to fix. They need you to treat an approval as a claim about today, not a receipt from last month.
If you have run a human approval step in production long enough to watch one go stale, that is what issue 348 needs next.
Top comments (0)