DEV Community

Anusha Mukka
Anusha Mukka

Posted on

Your Agent Reads Your Database. Your Attacker Writes to It.

SalesBleed showed that the scariest prompt injection does not arrive in a chat window. It arrives as a quiet row in your CRM.

Your lead form accepts strangers. That is the entire point of a lead form. On September 24, 2026, Zenity Labs published SalesBleed and showed what happens when one of those strangers writes a note your AI agent reads later as an instruction.

No employee clicked anything. No one typed at the agent. The attacker just filled out a form.

What SalesBleed Actually Was

Zenity Labs disclosed three vulnerabilities in Salesforce Agentforce on September 24, 2026. Salesforce had been told privately on June 1, 2026, and patched the reported bypasses in about two weeks. The timeline is worth noting because this is disclosure working the way it should: a researcher finds something serious, reports it quietly, the vendor fixes it.

Two of the three flaws enabled zero-click data exfiltration. Sensitive CRM data could be transmitted to attacker-controlled infrastructure without an employee clicking or approving anything. The third flaw let an attacker weaponize the trusted identity of an Agentforce-connected Slack agent to send phishing messages to employees from inside the organization.

The entry point was a feature nearly every Salesforce customer uses. Web-to-Lead lets external users submit data through a public form that flows directly into CRM records. It is designed to accept input from strangers. Zenity found that an attacker could plant hidden prompt-injection payloads inside those submissions. When a trusted Agentforce agent later processed the record, it read the attacker's instructions as if they were legitimate content and acted on them.

The exit door was Agentforce's Trusted URLs mechanism, an allowlist meant to control where the agent can send data. Zenity discovered the mechanism did not properly recognize top-level domains, and that certain character sequences could tamper with URL parsing, letting data slip past the allowlist to a destination the attacker controlled.

The model behaved exactly as designed. It read text and followed instructions. The flaw was assuming lead-form text was data, not instructions.

The Part Every Explainer Skips

Most prompt-injection writeups stage the fight in a chat box. An attacker types something clever, the agent obeys, everyone nods. SalesBleed's fight happened in storage. The payload sat in a database row, patient and ordinary-looking, until a trusted agent read it on an ordinary workday. The carrier was not a conversation. It was the CRM.

That changes the shape of the defense. If you picture prompt injection as a duel between an attacker and your agent, you build filters at the chat window and declare victory. But there is no chat window here. The instructions arrived through the same channel as every other lead, hours or days before they were read, and the only thing standing between the attacker and your data was the agent's ability to tell whose instructions it was following. A language model does not natively draw that line.

The durable lesson is not "filter harder." It is two things this piece will make concrete: mark where every field came from, and hold the exit door to a strict, correctly parsed standard. Most coverage of SalesBleed stopped at the injection. The interesting part is the exfiltration, because the allowlist failed at the parsing layer, and parsing layers are where your defenses quietly die.

Map the Lethal Trifecta Onto Your Own Agent

Security researchers describe the dangerous combination as the lethal trifecta, a framing from Simon Willison coined on June 16, 2025. An agent with private data access, exposure to untrusted content, and a way to communicate externally is exploitable by design. Any one of those legs is fine on its own. All three together mean a prompt injection buried in the untrusted content can turn the agent's own capabilities against its owner. Breaking any single leg blunts the attack.

I understand why all three are present, and that is worth saying before the scolding starts. A lead-enrichment agent without CRM access is a search box. Without the ability to reach external destinations, it cannot send the follow-up or the notification. Without ingesting leads, it has no job. The combination is the product. Nobody is going to remove a leg to make security reviewers comfortable, so the fix has to be boundaries around the combination, not subtraction of the combination.

Here is the attack path as a pipeline, because the order matters. The payload arrives first and detonates later, which is exactly why nobody noticed:

  [public lead form: anyone may write]
                |
                v
  +---------------------------+      +-----------------------------+
  | CRM record                |----->| trusted agent reads the row   |
  | attacker text sits here,  |      | "enrich this lead"            |
  | quiet and ordinary        |      +-----------------------------+
  +---------------------------+                |
                                               v
                               +-------------------------------+
                               | Trusted URLs allowlist        |
                               | mis-parses TLDs; character    |
                               | sequences tamper with parsing |
                               +-------------------------------+
                                               |
                       +-----------------------+------------------------+
                       |                                                |
                       v                                                v
            exfil to attacker-controlled                 Slack agent sends phishing
            infrastructure                               "from" a trusted identity
Enter fullscreen mode Exit fullscreen mode

Three things to notice in that diagram. First, the time gap between the top and the middle: the attacker does not need to be present when the agent acts. Second, the failure at the allowlist is a parsing failure, not a policy failure. The policy said the right thing. The code that enforced it could not tell evil-example.com apart from what it thought was allowed. Third, the rightmost branch costs the attacker nothing extra. Once the deputy is confused, every capability it holds becomes the attacker's capability, including the trusted Slack identity.

This is the confused-deputy problem wearing a new costume, a shape I walked through back in September with plugin supply chains. The deputy is trusted, connected, and authorized. The attacker never needed their own access. They only needed to get instructions in front of a deputy that could not tell whose instructions they were.

Tag Every Field With Its Author

The root cause is that models blur instructions and data, and no filter reliably unblurs them. The durable fix is to stop relying on the model's judgment and record provenance at the boundary: every field your agent reads should carry a tag saying who wrote it. Here is a small, runnable version in Python's standard library. It wraps the record your agent consumes and makes the trust boundary explicit instead of implied.

from dataclasses import dataclass
from urllib.parse import urlsplit

@dataclass(frozen=True)
class Field:
    """A value plus its provenance: who put it here."""
    value: str
    trusted: bool  # False if it came from outside your organization

def load_lead_record(raw: dict) -> dict[str, Field]:
    """Tag every field at ingestion. Untrusted until proven otherwise."""
    return {
        key: Field(value=str(value), trusted=False)
        for key, value in raw.items()
    }

def mark_verified(record: dict[str, Field], key: str) -> dict[str, Field]:
    """Only a human review or an internal system may flip a field to trusted."""
    record = dict(record)
    record[key] = Field(value=record[key].value, trusted=True)
    return record

def agent_reads(field: Field, action: str) -> str:
    """
    The tool boundary. Untrusted fields are data, never instructions.
    If a field contains instruction-like text and is untrusted,
    the agent reports it instead of obeying it.
    """
    lowered = field.value.lower()
    looks_like_instruction = any(
        phrase in lowered
        for phrase in ("ignore previous", "disregard", "send to",
                       "http://", "https://", "exfiltrate")
    )
    if not field.trusted and looks_like_instruction:
        return f"FLAGGED for review: untrusted field attempted '{action}'"
    return f"proceeding with '{action}' on trusted content"
Enter fullscreen mode Exit fullscreen mode

A few things are worth noting about this example. First, the default is untrusted. Fields do not earn trust by looking innocent; they earn it through a verified channel, a human review or an internal system. That single default is doing most of the security work. Second, the check lives at the tool boundary, not inside the model. The model can be as suggestible as it wants; the gate decides what counts as an instruction. Third, the flagging path reports instead of silently dropping, because a blocked attack you never hear about is a threat-intelligence feed you are throwing away.

Let me be honest about what this is. The keyword check is deliberately crude. In production you would want a classifier, an embedding-distance check, or a structured-output schema that never accepts free text where commands live. The point of the example is the shape: provenance tagging plus a boundary that refuses to promote data to instruction status. Swap the detection heuristic for whatever your team trusts. Keep the shape.

Hold the Exit Door

SalesBleed's second lesson is the one that gets less attention: the exfiltration succeeded because URL parsing, the enforcement layer of the allowlist, was wrong. Here is the thing about allowlists. Everyone agrees on the policy. Almost nobody audits the parser. If you take one mechanical task from this piece, make it this one: treat your egress URL check with the same suspicion you give your authentication code.

def egress_ok(url: str, allowlist: set[str]) -> bool:
    """Strict egress check. Exact hostname match, https only."""
    try:
        parts = urlsplit(url)
    except ValueError:
        return False
    if parts.scheme != "https":
        return False
    if parts.username or parts.password:
        # userinfo tricks: https://allowed.com@evil.com/
        return False
    host = (parts.hostname or "")
    if host != host.lower() or host.endswith("."):
        # demand canonical spelling: your own agent writes these URLs,
        # so it has no excuse for creative capitalization
        return False
    if host != host.encode("idna").decode("ascii"):
        # non-ASCII hostnames need explicit review, not silent acceptance
        return False
    return host in allowlist

# The part most teams skip: a test file of hostile URLs.
HOSTILE = [
    "https://allowed.com@evil.com/",          # userinfo trick
    "http://allowed.com/",                    # wrong scheme
    "https://allowed.com.evil.com/",          # subdomain of attacker
    "https://ALLOWED.COM./",                  # trailing dot, case games
    "https://xn--allowed-9nf.com/",           # lookalike punycode
]
SAFE = ["https://allowed.com/report"]

for url in HOSTILE:
    assert not egress_ok(url, {"allowed.com"}), f"leaked: {url}"
for url in SAFE:
    assert egress_ok(url, {"allowed.com"}), f"blocked legit: {url}"
print("egress gate holds")
Enter fullscreen mode Exit fullscreen mode

What to notice here: the check rejects by default on anything it does not understand. A parse failure is a deny, not a shrug. The userinfo rejection catches the classic allowed.com@evil.com construction. The canonical-spelling demand (exact lowercase, no trailing dot) removes the whole class of normalization mismatches instead of trying to enumerate them. The IDNA line forces punycode lookalikes into the open instead of silently normalizing them. And the test list is the real deliverable. An allowlist without a hostile test suite is a hope, and the SalesBleed Trusted URLs bypass is exactly what a hope looks like under adversarial pressure. Zenity beat the filters with obfuscated payloads elsewhere in the same disclosure cycle, which tells you the bar your tests need to clear: your own team, thinking like the attacker, before someone else does it for you.

Where This Breaks

No hedging on these. If you deploy the two gates above, here is what still gets through.

  • Provenance does not stop a trusted human from pasting hostile text. A compromised partner, a malicious employee, or a poisoned vendor feed all arrive wearing the trusted tag. Provenance is only as good as source hygiene.
  • A determined exfil can use destinations you allowed. DNS queries to an approved resolver, data hidden in error messages or image pixels, timing channels: an allowlist narrows the exit, it does not seal it.
  • The IDNA check above is a blunt instrument. Legitimate internationalized domains get flagged for manual review, which is a support ticket waiting to happen. Tune the rule to your actual traffic or accept the tickets.
  • Keyword detection at the tool boundary is a tripwire, not a wall. It catches the clumsy payloads. A patient attacker paraphrases.
  • Human approval on consequential actions kills the zero-click property, but approval fatigue turns the Approve button into a reflex. If your approvers click through, you have added latency, not security.

Build It If, Skip It If

Build it if your agent reads rows, tickets, emails, or documents written by people outside your trust boundary, and it can reach the network. That is the lethal trifecta wearing your company's logo, and SalesBleed is the receipt.

Skip it if your agent only answers questions over retrieved context and holds no tools, no write access, and no network path. Then your problem is answer quality, not exfiltration, and these gates are machinery without a threat.

The minimal viable version fits in an afternoon. One egress function with an exact-hostname allowlist, a test file of hostile URLs that you run in CI, and a log line on every blocked attempt. Watch those logs for a week. You will learn what your agent actually tries to reach, which is the cheapest threat model you will ever build. Provenance tagging comes second, after the logs have told you which fields are worth tagging.

Prototype the Gate, Then Read Your Own Logs

SalesBleed is worth your afternoon because it is not a Salesforce story. It is the textbook failure mode of every agent wired into real business data: untrusted content flowing toward a trusted deputy, with the exit door guarded by a parser nobody audited. The vendor patched quickly, the researcher disclosed responsibly, and the category remains wide open.

Build the egress gate. Feed it hostile URLs. Tag your fields at ingestion and refuse to promote data to instructions at the tool boundary. Then read the logs, measure what your agent actually attempts, and decide from the data whether you have a filter problem or an architecture problem.

What is the strangest destination your agent has ever tried to reach, and did anything stop it?

Top comments (0)