An agent that pays one invoice needs permission for that payment. Hand it an API key or a signing key and it has permission for everything the key can do, including whatever a prompt injection talks it into.
The alternative is to let the agent propose and have something else decide. The agent emits a structured intent. A policy engine checks it. A separate signer, holding keys the agent never sees, signs only what the policy approved. For high-value actions, a person confirms the exact transaction on a hardware screen.
Charles Guillemet, CTO of Ledger, laid out that design in a July 2026 podcast conversation with me and summed up the end state:
And the second one is security, like the keys stay in an enclave and the enclave doesn't sign anything if it's not approved.
Charles Guillemet, CTO of Ledger, on Chain of Thought ep 65
Your job as the developer is to make "approved" precise enough for a signer to enforce.
Two different ways it goes wrong
Guillemet separated two failure modes, with different fixes:
- The first is alignment: what you asked for and what the agent executes can drift apart. Natural language is ambiguous, and agents aren't deterministic.
Secondly, agents are not deterministic at all, there are plenty of heuristic, probability, so sometimes they will do exactly what you wanted, sometimes something slightly different.
Charles Guillemet, CTO of Ledger, on Chain of Thought ep 65
"Pay the supplier" leaves the destination and the amount open to interpretation, and no attacker is needed for the agent to pick the wrong one.
- The second is plain secrecy. To act on your data or money, the agent needs credentials: API keys, or in crypto the 24-word seed. If the agent holds them, an attacker can prompt-inject it into revealing them, or the agent can mistakenly reveal them. Charles pointed to a recent case of an account on X that was prompt-injected just by people tweeting at it.
A structured intent makes the proposed action something you can inspect. Keeping keys in a separate signer takes them out of reach of the model. You need both, because a well-protected key will still sign a bad transaction if the approval path lets it through.
Make the intent small and explicit
Start with one narrow operation, like a payment in a fixed currency. Define exactly the fields the policy needs and reject anything malformed before evaluating it. Amounts are integers in the currency's smallest unit. Recipients are exact identifiers from an allowlist. Expiry timestamps are timezone-aware. The agent never supplies the policy limits or the running total.
Illustrative sketch, not a production library or anyone's actual implementation. authorization_transaction and request_signature are hypothetical helpers for trusted state and the enclave-backed signer.
from dataclasses import dataclass
from datetime import datetime
@dataclass(frozen=True)
class Intent:
request_id: str
recipient: str
amount_minor: int
expires_at: datetime
def authorize(intent: Intent):
with authorization_transaction() as tx:
policy = tx.verified_policy
allowed = (
intent.recipient in policy.recipients
and 0 < intent.amount_minor <= policy.per_intent_limit
and tx.committed_today + intent.amount_minor
<= policy.daily_limit
and tx.now < intent.expires_at
)
if not allowed:
raise PermissionError("Intent outside policy")
approval = tx.reserve_once(intent)
return request_signature(intent, approval)
This runs in the trusted authorization service, not in the agent. reserve_once has to reserve the budget and record the request ID atomically, and committed_today has to count pending reservations as well as completed payments. Otherwise two concurrent requests can both pass the daily limit against the same balance. Retrying an identical request should return the original approval; reusing an ID with different contents should fail.
Bind the approval to what gets signed
The interface that matters most sits between the policy engine and the signer. An approval should name the exact intent and the policy version that allowed it, and the signer should authenticate that approval before it signs. An approved=True field the agent can set carries no authority.
Use a deterministic encoding of the approved transaction that covers every field that could change its effect, including ones this sketch fixes outside the intent, like currency and network. The signer rejects anything that differs from the approved contents and checks expiry again at signing time, since a request can go stale while it waits for a human.
Guillemet wants guarantees on the policy engine too. He described two ways to get them: run it inside a secure enclave, or prove it ran correctly with a zero-knowledge proof. He was candid about the first option: "TEE is not that secure, but it's better than full software." Moving a Python function into another process gives you none of those guarantees, so treat execution integrity as a deployment requirement, separate from the code above.
Policy changes need their own path. If the agent can add a recipient or raise a limit, it can loosen its own constraints. His suggestion is to review and sign the policy itself with a hardware device, so the policy engine knows the policy came from you. Mind you, as CTO of Ledger, there might be a bit of bias there - but it's certainly not a bad idea.
Match the friction to the asset
Charles was clear that there's no single right answer. It depends on your threat model and on classifying what you're protecting. He sketched out four levels:
| What you're protecting | Approach | What it buys you |
|---|---|---|
| Low-value assets | Software and access control | May be enough when the loss is tolerable |
| Something with more value | Hardware that signs automatically (a HashiCorp Vault-style setup) | Keys stay protected, but execution isn't verified |
| Higher-stakes actions | Hardware with a screen | A person sees and confirms what gets signed |
| The most critical actions | Secure display plus multisignature | A quorum has to approve |
Ledger uses that top level on itself for firmware releases:
We are using devices with a screen where we can verify what we sign in a multi-signature setup.
Charles Guillemet, CTO of Ledger, on Chain of Thought ep 65
For your own system, decide which agent actions need on-device confirmation. The screen should show the destination and amount of the transaction being signed. A generic "Approve?" prompt tells the reviewer nothing, and any edit after approval should require approving again.
Takeaway
Before you give an agent payment authority, check that:
- the agent submits a structured intent and never holds the signing key;
- trusted policy enforces recipients and limits, and the agent can't edit the policy;
- accounting handles concurrent requests and retries;
- the signer verifies an approval bound to the exact transaction and rechecks expiry;
- high-value actions need on-device review, with multisig where the threat model calls for it.
The full conversation, with the transcript, is on Chain of Thought. You can also watch it on YouTube.
Drafted with AI assistance from the episode transcripts, then edited by me.
Top comments (0)