If you write backend code, you know the SQL injection fix by heart. Here’s the same bug twice, then what carries over to LLM features and what doesn’t.
Same bug, two languages
1- # Vulnerable: input becomes part of the command
cur.execute(f"SELECT status FROM orders WHERE email = '{email}'")
Fixed: command and data travel separately
cur.execute("SELECT status FROM orders WHERE email = %s", (email,))
2- # Same shape. No placeholder syntax exists for this one.
prompt = f"""You are a support assistant. Only discuss order status.
Customer message:
{user_message}
Knowledge base excerpts:
{retrieved_chunks}
"""
reply = llm.generate(prompt)
There’s no bind parameter for the second block. Chat APIs separate system and user roles, and that helps the model prioritize, but it’s trained behavior, not an enforced boundary. The UK NCSC’s December 2025 write-up says the same thing from the other direction: no inherent distinction between data and instruction, so prompt injection may never be fully mitigated the way SQL injection can be.
What carries over from the SQLi playbook
3 habits transfer directly. Give the “database user” (here, the tools) the minimum rights. Treat everything the model emits as untrusted. And never derive authorization from a claim the model makes.
from pydantic import BaseModel, Field
class OrderStatusArgs(BaseModel):
order_id: str = Field(pattern=r"^[A-Z0-9]{8}$")
ALLOWED_TOOLS = {"get_order_status"} # read-only
def run_tool(call, session):
if call.name not in ALLOWED_TOOLS:
raise PermissionError(f"tool not allowed: {call.name}")
args = OrderStatusArgs.model_validate(call.arguments)
# identity comes from the session, never from model output
return get_order_status(
customer_id=session.customer_id,
order_id=args.order_id,
)
OWASP’s advised mitigations point the same way: limit which tools an agent can use, what each tool can do, and what permissions it holds.
A cheap tripwire (not a fix)
import secrets
CANARY = f"cnry-{secrets.token_hex(6)}" # generate per deployment, keep in config
SYSTEM_PROMPT = f"...internal ref {CANARY}..."
def check_output(text: str) -> str:
if CANARY in text:
alert("system prompt leak") # log it and block the response
return "Sorry, I can't help with that."
return text
This only catches verbatim leaks. An attacker who asks for a paraphrase or an encoding walks right past it. Use it as detection, never as protection.
Test it like the attacker gets to iterate
In NIST’s agent hijacking tests, the strongest standard attack succeeded 11% of the time, while attacks tailored to the target hit 81%. In “The Attacker Moves Second,” researchers from OpenAI, Anthropic, and Google DeepMind broke 12 published defenses with adaptive attacks, most above 90% success, though most had reported near-zero originally. Each defense used its own benchmark, so read the numbers as directional.
A static payload list in CI is fine as a regression check. It isn’t evidence your system is safe.
A design budget- the Rule of 2
Meta’s Agents Rule of 2 says that within one session an agent should satisfy no more than 2 of 3 properties: processing untrusted input, accessing sensitive systems or private data, and changing state or communicating externally. If a feature needs all three, that’s the one to put a human approval step in front of.
Top comments (0)