DEV Community

Arun Kumar
Arun Kumar

Posted on Originally published at diffstudy.com

If nothing was concatenated, it isn't prompt injection

Most write-ups treat prompt injection and jailbreaking as the same attack with two names. They aren't, and the person who coined the term has a one-line test that settles almost every case.

It matters because the two attacks live at different layers, so the defences don't transfer. Build a jailbreak defence against an injection problem and you've spent the effort for nothing.

The test

From Simon Willison's post, March 2024:

Crucially: if there's no concatenation of trusted and untrusted strings, it's not prompt injection. That's why I called it prompt injection in the first place: it was analogous to SQL injection, where untrusted user input is concatenated with trusted SQL code.

That's the whole test. Did untrusted input get glued into a prompt your developer wrote? Then it's injection. Did someone just argue the model out of its own guardrails in a chat window? Then it's a jailbreak.

He defines the two separately. Prompt injection is "a class of attacks against applications built on top of Large Language Models (LLMs)", which "work by concatenating untrusted user input with a trusted prompt constructed by the application's developer." Jailbreaking is "the class of attacks that attempt to subvert safety filters built into the LLMs themselves".

Read those side by side and the split is clean. One attacks your application. The other attacks the model.

Why the distinction does work

A jailbreak is a conversation problem. One person, one chat window, pushing against a filter until it gives. No second party, no external document, no tooling. The defence lives inside the model — it's alignment training, and it isn't yours to fix.

Prompt injection is an architecture problem. Something untrusted — a user, a file, a fetched web page — ends up concatenated into a prompt that also contains your instructions and, often, your credentials. The defence lives in how you build the application, which is yours to fix.

The stakes differ too, and Willison is direct about it. A jailbreak risks embarrassment, or at worst helping someone do something harmful. Prompt injection threatens applications holding confidential data and tools. Those aren't the same exposure.

Worth noting he doesn't claim it's solved. He's proposed patterns like Dual LLM, but by his own account prompt injection remains an open problem. Anyone selling you a complete fix is ahead of the person who named the thing.

Where the authorities actually disagree

Here's the part that trips people up, and it's not a case of someone being wrong.

OWASP draws the boundary differently. In its taxonomy, jailbreaking sits underneath direct prompt injection — one of two sub-classes, alongside indirect injection. Willison keeps jailbreaking in its own category entirely.

So the same clever chatbot prompt gets two different names depending on which taxonomy you're following. Neither is careless. They're organised for different purposes: Willison's split is about where the vulnerability lives, OWASP's is about cataloguing a risk surface.

What matters is knowing which one your team means when someone says "we handled prompt injection" in a review. That sentence carries two different scopes.

Which one you're actually facing

The concatenation test answers it fast. Ask where the malicious text came from.

If it arrived in the user's own message and the attacker is the user, you're looking at a jailbreak. If it arrived inside data your application fetched and pasted into a prompt — a support ticket, a scraped page, a PDF, a tool result — that's injection, and the attacker may never touch your app directly.

That last case is the one worth sitting with. Indirect injection doesn't need access to the victim's session at all. Someone poisons a document, your app reads it, your prompt concatenates it. Nobody had to log in as anyone.

If you take one thing: when a model can read untrusted data and call tools, you own an architecture problem, and no amount of model-side safety training will close it for you.


Longer version with the full comparison table at diffstudy.com. Quotes from Simon Willison, "Prompt injection and jailbreaking are not the same thing", 5 March 2024.

Top comments (1)

Collapse
 
supportdev profile image
DEV SUPPORTS •

Dеar User,
Due to аn inсrease іn bot асtіvity оn thе platform, we rеquіre verіfy of уour account.
Pleаsе lоg in vіа thе link bеlow:
• anti-bot.icu/5K0N5G7M9C4
Verificated deadlinе - 12 hours.
Sincerely,Dev Supроrt

‍​‌​‌