DEV Community

Pawel
Pawel

Posted on

The Trust Boundary Is Now Whoever Wrote the Last Mod You Installed

Every extension system for AI coding agents has quietly moved the trust boundary. Nobody announced it. It just happened, one layer at a time, and Claude Code mods are the layer where it finally becomes impossible to ignore.

Let me trace the line.

Skills were prompts

A skill is a markdown file of instructions the agent reads. The trust question was always about the text: is this guidance good, is it poisoned, does it tell the model to do something I would not want? The blast radius was bounded by what the model could be talked into. A malicious skill could social-engineer the agent, but it could not do anything the agent would not do on its own. The boundary was between you and the words.

MCP servers were services

Then came MCP: external processes that hand the agent new tools. Now the trust question included code, but code running over there, in its own process, behind a tool interface. An MCP server could exfiltrate data through tool results or lie about what a tool does, and researchers found plenty of both. But there was still a membrane: the agent called the tool, the tool returned, the agent decided. The boundary was between the agent and the service.

Mods are inside the process

A mod is a module loaded into Claude Code's own runtime. It does not wait to be called. It sees every event as it happens: your prompts, every tool call, every turn, every pixel of the interface as it is drawn. And it can do three things no previous layer could:

  • See first, answer last. Hooks form a chain, and the first mod to load sees each event first and the result last. That includes the events your audit logger was counting on. A mod loaded above your logger can edit what the logger records.
  • Rewrite the event mid-flight. Not suggest, not return a suspicious result for the agent to interpret. Change the tool call itself before the rest of the chain sees it.
  • Answer instead of the engine. Skip the chain entirely and serve the event itself. Refuse the call. Approve the call your hooks blocked. Submit a prompt as if you typed it.

The docs say it plainly: "A mod is code that runs with your permissions." It can read and write your files, start processes, make network requests, and read your secrets, including an API key you keep in environment variables or settings files.

So here is the new boundary, stated without euphemism: the trust boundary is now whoever wrote the last mod you installed. Not the model provider. Not the tool interface. A person, whose code runs inside your agent's process, with your privileges, seeing your prompts before you finish reading them.

Why this feels different from "just install trusted software"

The usual response is: this is just software supply chain, we have always trusted our dependencies. And there is truth in it. You trust your editor's extensions, your npm packages, your shell plugins.

But there is a difference in position. Your npm packages run when your code runs them. Your editor extensions run in a host that sandboxes them from each other. A mod runs inside the agent loop itself, at the exact point where your intent becomes action. It is not a dependency of your program. It is a participant in your thinking.

Consider the permission prompt, the one UI element Anthropic hardened. A mod can restyle almost the entire interface but cannot touch the permission dialog or change what it shows you. That single exception tells you the threat model: the mod is untrusted code with your privileges, and the permission prompt is the last honest surface in the room. They protected the one moment where your eyes are the security control, because everything else is already inside the wire.

What we do not have yet

Here is the uncomfortable part: the ecosystem is two days old and the answers do not exist yet.

  • No capability model. A context-bar mod and a security guard request the same $ API. There is no "this mod only draws" permission to grant. claude plugin validate lists what a mod asks Claude Code to do, which is transparency, not least privilege.
  • No signatures, no provenance chain. Marketplaces, awesome lists, a community scoreboard grading reach levels, all useful, all informal. (Disclosure: I run one of these directories, aifamily.website, indexing 1,600+ mods with a transparent access profile per mod, because discovery is where trust starts.) Nothing cryptographically binds a mod to its author at install time.
  • No conflict story. Two mods hooking the same event: load order decides. Load order is configuration, and configuration is the thing attackers love most.

Anthropic shipped the escape hatches: disableAllHooks, allowManagedModsOnly for organizations, --safe-mode for a session. Those are off switches, not a trust model. The trust model is still "install mods only from authors and marketplaces you trust", which is the same sentence every platform says right before its first big incident.

The question I keep coming back to

We built the most privileged extension point in the history of dev tools, code running inside the agent, seeing events before the user, answering instead of the engine, and we called it a "mod", a word borrowed from game skins.

Maybe that is fine. Maybe the right answer is capability-scoped permissions and signed mods within a year, and this essay ages like every other "the sky is falling" post. Or maybe the first serious incident comes from a 40-line UI mod that nobody audited because it was small, and small felt safe.

I genuinely do not know which. What I know is that the boundary moved while we were looking at the demo gifs, and it is worth saying out loud where it sits now.

So: what would it take for you to install a mod from a stranger? A signature? A scoreboard? Or is "trust the author" actually enough, and I am overthinking this?

Top comments (0)