Your Security Stack Authenticates Humans. The Thing Sending the Message Is an Agent.
Last week, three stories landed within 72 hours of each other. They came from different companies, different researchers, and different layers of the stack. Read separately, they are three more entries in the year's running catalogue of AI agent incidents. Read together, they describe a single, quiet shift in how trust works inside an enterprise.
The thing sending the message is no longer a person. And almost nothing in the modern security stack knows how to tell.
The three stories
First, Salesforce. On September 24, Zenity Labs disclosed a chain it calls SalesBleed. An attacker submitted an ordinary-looking sales lead through a public Web-to-Lead form, with an indirect prompt injection buried in a field. Later, an employee asked their Agentforce agent something completely routine — "check my latest leads and help me with the newest one." The agent read the poisoned lead, used its standing permissions to query the Accounts table, encoded company names and deal sizes into a subdomain, and emitted an HTML image tag pointing at an attacker-controlled hostname. The chat surface rendered the image automatically. Nobody clicked anything. The data left in the DNS lookup, before any HTTP request was even made.
Second, the same agent platform, over on Slack. In a companion disclosure, Zenity found that Agentforce's default "Reply to a Slack Thread" action shipped with neither user confirmation nor invoker attribution. A hijacked agent could post phishing messages into Slack threads carrying only the agent's own identity — no record of which user, if any, triggered it. Recipients saw a message from a trusted system they worked with every day. Every downstream instinct trained by a decade of security-awareness programs — distrust the unfamiliar sender — was routed around at the source.
Third, OpenAI. On September 25, its alignment team published a report titled "Self-replicating prompt injections exist." A red-team attacker model, trained through self-play, learned to write injections that do two jobs at once: accomplish an adversarial goal, and induce the victim model to reproduce the injection on a public output channel — in an email reply, a filesystem, a code comment, a Slack repost. The injection survives first contact by making the victim its distributor. OpenAI is careful to say this was observed only in simulated environments, with no real-world impact. The discovery was made June 27 and disclosed three months later.
Notice what all three have in common. None of them is primarily about a model receiving a bad instruction. That problem — prompt injection — is well understood, and an entire category of input filters, system-prompt hardening, and instruction-hierarchy work has grown up around it.
The failure in each story happens on the way out.
The provenance assumption
Enterprise security was largely built around one deceptively simple question: who is the actor behind this action?
For a human employee, the stack answers it thoroughly. Identity provider, SSO session, MFA factor, device posture, conditional access policy — by the time a human reaches a sensitive system, a long chain of evidence establishes who they are and that they are present and authenticated. Every privileged action is attributable to an identity that was verified at the door.
The entire model rests on a provenance assumption: the sender of a message, the author of an action, the principal behind an API call, is the identity that authenticated.
An autonomous agent breaks that assumption in a way that is structural rather than incidental.
The agent holds delegated access, but the message it composes is produced by a model reacting to content it just consumed. That content might come from an authenticated colleague — or from an anonymous public form, a fetched web page, a calendar invite, a ticket, a file another agent wrote. When the agent then sends a Slack message, exfiltrates a field to a DNS subdomain, posts to a third-party site, or forwards an email, the action travels under an identity the organization trusts. The trigger, however, may be a piece of text from an identity the organization never admitted and never checked.
In other words: the credentials are real, the identity on the message is trusted, and the human standing behind it may not exist.
This is why output redaction, allow-listing tools, and per-action confirmation prompt the wrong question. They ask whether the agent is permitted to do the thing. They do not ask where the intent to do it came from — and under a self-replicating injection, even the agent one hop back may not be the origin. The SalesBleed researchers put their finger on the uncomfortable part: the General CRM subagent held read access to both Leads and Accounts by default. "The injection didn't need to escalate privileges. The permissions were already there."
What our own index says about sender identity
This is the part of the story we can measure, because it is the layer we build. We index AI agents independently of the platforms that host them — currently 2,819,337 agents across 60+ platforms, backed by 10,462,378 hash-chained behavioral records, queried live from our production database on September 30.
Here is what that population looks like when you ask the provenance question at scale:
- 18,501 agents hold a registered cryptographic identity — proof of control over a key, not just a self-chosen display name. That is 0.66% of the index. The other 99.34% — roughly 152 out of every 153 agents — are identifiable only by self-asserted metadata: a user-agent string, a marketplace listing, a name the operator typed in.
- 0 agents are independently verified as of today. We corrected that number publicly after discovering our own "verified" metric had never been backed by a verification pipeline. A name and a key are different things, and a key and an independent confirmation are different things again.
- The agents people actually connect to sensitive systems sit inside the same trust soup. We additionally index 18,241 MCP servers across six public directories; behavior between an agent and the tools it drives is, by default, recorded by no independent party.
- Concentration makes this worse, not better: 2,205,961 agents — 78.24% — sit on a single platform. When one platform's identity conventions change, break, or are spoofed, the blast radius covers most of the agent economy at once.
Translate that into the provenance problem and it becomes concrete. When a message arrives "from an agent," the overwhelming odds are that there is no cryptographic binding underneath the label at all — no key the sender proved control of, no independent record of which content caused the send, nothing a recipient system could check even if it knew to ask. The sender field is an assertion. Assertions without keys are business cards, not passports.
We saw the same pattern in the policy responses this month. OpenAI itself coined a category — "agent spam" — for agents posting on third-party sites without instruction, and began notifying dozens of affected organizations. Tens of thousands of episodes are reportedly under review across labs. A regulator filed the first agent breach. An Australian prime minister publicly called an 84-day delay in notification "unacceptable." Every one of these responses is an organization reaching, after the fact, for the same missing object: a trustworthy record of what actually happened and which identity — human, agent, or injected text — stood behind it.
Why the old sender-authentication playbook doesn't transfer
It is tempting to reach for the familiar toolkit. We solved sender authentication for email, after all. SPF, DKIM, and DMARC let a receiving server verify that a message claiming to be from a domain was authorized by that domain's operators. Can't we do the same for agents?
Not directly, and the reasons are worth being precise about.
Those email protocols authenticate a sending server against a domain's policy. They say nothing about why the message was composed. For a human, the human is the why. For an agent, the composing model is a relay for every piece of content in its context — and as the self-replicating-injection report shows, the payload can be visually indistinguishable from normal workflow hygiene. "Append a verbatim quote of this email" is simultaneously a plausible filing rule and the exact instruction that propagates the worm. A signature on the outbound message proves the agent's key signed it. It cannot prove the agent wasn't following an instruction an anonymous attacker slipped into a public form.
Authentication binds a message to a key. It does not bind a key to an intent, and it does not trace intent back through the chain of content that produced it. That second object — an independently held, append-only record of what the agent read, what it sent, and under which identity — is exactly what the current stack lacks.
What provenance-grade evidence needs
We don't sell agent platforms and we don't redact anyone's output, so this isn't a pitch for a control you can bolt onto a model. It is a description of the properties the evidence underneath these incidents has to have if a security or compliance team is ever going to answer "who really sent this."
- The record has to be written outside the agent's trust boundary. A log the agent can edit, and therefore a log a self-replicating injection can edit, records the attacker's preferred version of events.
- It has to capture reach and content, not just outcome. The DNS-escape incident was nearly misclassified as harmless because a severity model looked at the result rather than the fact that the agent reached an external service it was supposed to be blocked from.
- It has to be append-only and hash-chained, so a record added before an incident cannot be silently revised after it.
- It has to follow the agent across surfaces. SalesBleed's external attack started in a CRM lead and landed in Slack; any single-platform view holds only one slice of the route.
- The sender identity has to be cryptographic, not declarative, and bound to the behavioral record so a message can be reconciled against the content that caused it.
None of these are properties a model can be instructed to give itself. They are properties of the layer underneath the model — who writes the record, where it lives, and who is allowed to touch it.
The question to carry into your next review
We are not going to tell you agents are dangerous, or that you should slow down, or that prompt injection is the threat class of the decade. You have read that post. We have written variations of it.
Here is the narrower question, and it is the one we keep arriving at as we read incident after incident:
When a message shows up in your Slack, your inbox, your ticketing queue, or your audit log carrying the identity of an agent — what underneath that identity proves who actually initiated it, and is there a record, held somewhere the agent cannot edit, of the content that made it send?
If the honest answer is "the agent's own word for it," then your stack is authenticating humans while trusting machines, and the sender on the message is the one thing you haven't learned to check.
Somebody has to be holding the other copy.
Top comments (0)