DEV Community

CAI
CAI

Posted on

Why CAI exists: the agent-payment problem and the three roles (custodian, approver, operator)

Why CAI exists: the agent-payment problem and the three roles (custodian, approver, operator)

The agent-payment problem is older than agents, and it has a clear shape. This post is the H1-readable essay on the problem CAI solves, the user-confirmation pattern, and the three roles that make the system work. No code, no commands. This is the post for the reader who wants to understand the why before the how.

The problem

An agent wants to do something that costs money. The user is on the hook for the cost. The agent has no wallet, the user has no way to give it one without either pasting a private key (which the user should never do) or building a custodial flow from scratch (which most teams don't). The agent either fails at the task or asks the user to do it manually. The user pays anyway, but the agent did not help.

The same problem exists for site credentials. The agent wants to log in to a site on the user's behalf. The user has a password, the agent has no way to retrieve it without either the user pasting the password into the agent conversation (which the user should never do) or building a vault from scratch. The agent either fails the task or asks the user to do it manually.

Both problems have the same shape. The agent has the intent to do something the user has authorized, but the agent has no way to execute the action without the user giving up something they should never give up.

The historical failure modes

There have been three attempts to solve this problem. All three have well-known failure modes.

Attempt 1: the agent holds the user's private key. The user gives the agent the private key to their wallet. The agent signs transactions autonomously. The user has no visibility into what the agent signed.

Failure mode: a compromised agent drains the wallet. The user finds out after the fact. The damage is done.

Attempt 2: the agent holds the user's password. Same as attempt 1, but for site credentials. The user gives the agent the password to their bank, their email, their social media. The agent logs in autonomously.

Failure mode: a compromised agent logs in to every site the user has saved. The user's accounts are compromised.

Attempt 3: the user pastes the secret into the agent conversation. The user pastes the private key or the password into the chat. The agent's model sees the secret. The model may log it. The agent uses the secret for the action.

Failure mode: the secret ends up in the model's context window. The model may retain it. The user has to rotate the secret.

All three attempts fail for the same reason. The user gives up something they should never give up, and the agent's compromise becomes the user's compromise.

The three roles

CAI splits the responsibility into three roles.

The custodian. CAI holds the private key. CAI holds the encrypted credential. CAI signs the transaction when the user has consented.

The approver. The user sees the recipient, the amount, the chain, and taps once on the hosted-action page. Without the tap, nothing happens.

The operator. The agent calls the CAI API, returns the hosted-action URL, and polls for the receipt. The agent does not hold the private key or the password.

The agent's API key is the authorization. It can call the CAI API on the user's behalf. The user's tap on the hosted-action page is the consent. It authorizes a specific transfer.

This is the only configuration where the agent can do useful work without the user giving up something they should never give up.

The user-confirmation pattern

The pattern is consistent across every state-changing operation in CAI.

  1. The agent calls the CAI API with the operation details.
  2. CAI generates a hosted-action URL. This is a single-tap confirmation page.
  3. The user opens the page, sees the operation details, and taps once.
  4. CAI executes the operation.
  5. The agent polls for the receipt.

For a transfer, the page shows the recipient, the amount, and the chain. For a credential retrieval, the page shows the site and the username. The user sees the actual operation before it executes.

The page is HTTPS, the tap is bound to a single operation, and the URL expires in a short window. If the user does nothing, the operation does not execute.

Why this design works

The user-confirmation pattern works because it gives the user visibility and consent.

Visibility. The user sees the actual operation before it executes. The user can see the recipient's name or address, the amount, and the chain. The user can see the actual value leaving their wallet.

Consent. The user explicitly approves the operation. The tap is the consent. Without the tap, nothing happens.

The same pattern powers the safety limits. Every wallet has a default USD 200 per day auto-limit on agent-initiated transfers. Every transfer to a new recipient requires explicit user confirmation. These are product-level guards, not optional settings.

A compromised agent cannot drain the wallet. The agent can call the CAI API, but the API returns a hosted-action URL. The user does not tap. Nothing happens. The agent has been compromised, but the user is not at risk.

A compromised agent cannot log in to a site. The agent can call the vault API to retrieve a credential, but the API returns a hosted-action URL. The user does not tap. Nothing happens. The user is not at risk.

A compromised model cannot exfiltrate the user's secrets. The model sees the API call, the URL, and the user-confirmation page. It does not see the private key or the password. The secret never enters the model's context window.

The user-confirmation pattern is the practical implementation of "the user is in the loop on every privileged action."

The agent's job

The agent's job is to do useful work for the user, not to hold the user's secrets. The agent calls the CAI API, returns the hosted-action URL, and polls for the receipt. The agent never has the private key or the password. The agent's storage is the API key and the response payloads, both of which are scoped and time-limited.

This is a different model from "the agent is the user's digital self." CAI's model is "the agent is the user's tool." The user has the secrets. CAI holds them. The agent operates on them. The user approves every privileged action.

Why this matters for the agent ecosystem

The agent ecosystem has a trust problem. The user's wallet, the user's site credentials, and the user's identity are the things the user values most. They are also the things the user is most reluctant to give an agent access to. Without a way to give the agent access without giving up the secret, the agent cannot do useful work.

CAI is one answer to that trust problem. The user-confirmation pattern, the three roles, and the single-tap hosted-action page are the mechanisms that make the trust work. The agent ecosystem will need more of these mechanisms over time, and CAI is one of the first to ship.

The full system is documented at cai.com. The signup is at cai.com/app. The contract is at cai.com/skill.md.


If you tried the user-confirmation pattern described above and the hosted-action page did not render, or the user's tap did not authorize the operation

Comment below with:

  1. What you ran -- the install command, the request, the MCP host config. Copy the actual command or request.
  2. What you expected -- one sentence.
  3. What you got -- the error message, the empty response, the unexpected behavior. Paste it verbatim.
  4. Your environment -- OS, Node version, the MCP host (OpenClaw / Hermes / Codex / Cursor / other), the CAI account tier if relevant.

Every comment on this article gets read. Bug reports will be replied to within 24 hours. Friction points shape what we document next.

Documentation: cai.com/skill.md · cai.com/developers.html · cai.com/app to sign up.

Top comments (0)