DEV Community

WUMBOLOVER
WUMBOLOVER

Posted on Originally published at neverlosebitcoin.com

Your AI Agent Has a Bitcoin Wallet. Here Is the Threat Model.

If you are building AI agents with MCP, there is a good chance your agent can already move money. An MCP server can expose pay_invoice as easily as search_files, and to the model they are all just function calls. That symmetry is the whole security problem.

I just published a full deep dive on this at Never Lose Bitcoin: When AI Agents Hold Bitcoin: MCP Wallets and the Security Playbook. It covers every confirmed incident, the attack classes researchers have documented, and the defense checklist. Here is the developer summary.

How agents hold bitcoin in practice

Three patterns dominate:

NWC connection strings with budgets. Nostr Wallet Connect (NIP-47) gives the agent a scoped connection URI with a per-connection spending limit and instant revocation. Alby Hub, Coinos, Mutiny, and Zeus support it. One agent, one connection, one budget. Compromised agent? Revoke the string.

Pay-per-call. L402 turns a Lightning payment into API authentication: the agent pays a BOLT11 invoice and gets a credential from the preimage. No standing keys to steal. Coinbase's x402 does the same over USDC.

Purpose-built agent wallets. LNbits has an Agent Wallet extension with per-agent profiles, single-payment and daily limits, dry-run requirements, and approval thresholds. Lightning Labs' Wavelength (July 2026) exposes a self-custodial wallet over MCP while keeping seed operations outside the agent channel.

The incident you should know

In May 2026, Bankrbot lost roughly $200K in DRB tokens because an attacker replied to Grok in Morse code. Grok decoded it and reposted it as a plain-text instruction. Bankrbot treated it as authoritative and transferred 3 billion DRB. The key was never stolen. The custodian signed because the agent asked. No human was in the loop.

That is the shape of every agent-wallet incident: nobody breaks the cryptography, the model just gets talked into spending.

The attack classes

Researchers have documented the mechanisms even where thefts are not yet public:

  • Tool poisoning (Invariant Labs): malicious instructions hidden in tool descriptions, which the model reads to decide how tools work.
  • Line jumping (Trail of Bits): descriptions enter model context at listing time, before approval prompts appear at call time. The approval dialog becomes a rubber stamp.
  • Tool shadowing / name squatting: a malicious server registers the same tool name as a trusted one (CVE-2026-30856).
  • Confused deputy: a wallet MCP server holding an admin macaroon cannot distinguish a legitimate agent request from one the agent was tricked into making.
  • Supply chain: malicious npm/PyPI packages installing rogue MCP servers, including campaigns attributed to North Korea's Famous Chollima group.

The one design rule

Enforce spending policy where the key is used, never in the prompt. A limit written in the system prompt ("never spend more than $10") is a suggestion the model reads next to attacker instructions, and attacker instructions can override it. A limit enforced by wallet infrastructure code, which never reads model output, cannot be talked out of existence.

Practically: NWC budgets per agent, human approval above a threshold (LNbits supports this, ln-mcp fires a webhook), one wallet per agent funded like a hot wallet, pinned and hash-verified MCP servers, full invocation logging, and a one-click revocation path. For larger amounts, put the agent behind multisig so it can propose but never dispose alone.

The full deep dive has the complete incident timelines, the defense playbook in checklist form, and a concrete safe-stack example: When AI Agents Hold Bitcoin: MCP Wallets and the Security Playbook.

Top comments (0)