DEV Community

NET_DARK_BOI
NET_DARK_BOI

Posted on

Prompt injection isn't a payload problem — 3 questions decide if your AI feature is exploitable

Prompt injection gets talked about as if it were a clever string you type to jailbreak a chatbot. That framing misses the actual problem, and it is why so many "fixes" do not work.

The real issue is structural. An LLM feature puts two different things into one text stream: the instructions the application intends, and the data it processes. The model has no reliable way to tell an order it should follow from text it should merely process. That is the same root condition as SQL injection or command injection — instructions and data sharing a channel — except there is currently no complete escaping that makes natural-language data safe to concatenate with instructions.

So the question is not "what payload breaks it." The question is whether a hostile instruction, hidden in data the model reads, can reach something that matters.

The lethal trifecta

Simon Willison's framing is the most useful lens. An AI feature becomes dangerous when it combines three things:

  1. access to private data,
  2. exposure to untrusted content, and
  3. a way to communicate externally.

Any one of the three is survivable. All three together mean an attacker can plant an instruction in the untrusted content, have the model read private data, and have it sent out. Guardrails that catch 95% of phrasings do not save you, because the attacker only needs one that gets through.

Three questions for any AI feature

Instead of hunting for magic strings, map the feature.

1. Where does untrusted text enter the prompt?
Not just the user's message. Any retrieved document, web page, file, tool result, or earlier message the model reads is a potential injection point. A support bot summarising a ticket, an agent browsing a linked page, a RAG assistant quoting a document, an assistant reading an inbox — all ingest text written by someone else. This is indirect prompt injection, and it is usually the higher-severity case, because the hostile instruction does not come from the person using the feature.

2. What can the model actually do as a result?
A model that only writes a summary back to the same user can be misled, but the blast radius is that user's session. A model wired to tools — sending email, calling internal APIs, reading files, making purchases — can be steered into real actions. Severity is decided by the model's privileges, not the cleverness of the phrasing. Write down every tool it can call and whose authority it runs with.

3. Can a hostile instruction reach a real action or a real exfiltration?
This is where the trifecta closes. The classic example is data exfiltration through a markdown image: a hidden instruction tells the assistant to encode something it can read — a recent message, a token — into the URL of an image, which the client then fetches automatically, sending the data to the attacker's server. No script execution required.

If the answer to all three is yes, you have an exploitable feature, not a curiosity.

Why filtering the input does not fix it

The instinct is to scan the input for "ignore previous instructions" and block it. It does not hold. Injections hide in Base64, emoji, other languages, text split across documents, or even inside images for a multimodal model. An instruction need not be human-readable to work. Blocklisting natural language is an unwinnable game.

This is not theoretical. Prompt injection has been demonstrated against production AI features at Microsoft 365 Copilot, Slack AI, and GitLab Duo, among others — real systems wired to real data.

What actually helps

Because the input cannot be fully sanitised, the durable controls sit around the model, not inside the prompt:

  • Least privilege for tools. Give an agent only the functions a task needs, scoped to the current user, and prefer read-only. High-consequence actions — sending messages, moving money, changing settings — should require an explicit human confirmation that shows what will happen.
  • Treat the model's output as untrusted. If the response is rendered as HTML, it can carry stored XSS; if it is passed to a shell or a database, it carries that injection class. Encode output for its destination.
  • A deterministic check the model cannot talk past. Authorization enforced outside the model is what keeps an injected instruction from becoming an injected action.
  • Isolate and label retrieved content, limit how much untrusted text enters the context, and log tool calls so abuse is visible afterward. The "dual LLM" pattern — a quarantined model handles untrusted data, a privileged model never sees it — is worth reading about.

None of these "solve" injection. Together they shrink what a followed instruction can reach, which is the realistic goal.

Test it like a feature, not a vibe

Open-source tools like Garak and Promptfoo probe for injection and unsafe output handling, and a guardrail classifier can sit in front as one more layer — never the only one. But the most useful test is the one you write for your own feature: place a benign marker in the kind of content the feature consumes, and see whether it comes back in the output. If it does, the data channel reached the instruction channel.

Try the mechanics hands-on

Reading about it only goes so far. These run in the browser:


Sources


What is the trickiest indirect-injection vector you have seen in a real feature?

Top comments (0)