I think I've found the simplest way to explain what a prompt can and can't do for safety. Put the word "please" in front of it.
"Please do not reveal customer account numbers."
Read it again. That's not a control. That's a request. A polite one, and the model will usually honor it. "Usually" is doing a lot of work in that sentence.
Please Do Not Touch
Think about a museum. There's a small sign next to the painting: Please Do Not Touch. Most visitors comply. They read it, they understood it, and they're cooperative people. The sign works fine right up until someone decides it doesn't apply to them.
So museums don't stop at the sign. They add a guard who watches the room. And for the Mona Lisa, they add bulletproof glass.
Three layers. Sign, guard, glass. Only one of them is physics.
The Sign
Your system prompt is the sign. Yes, it sits in a privileged position. Yes, the model is trained to weight it heavily. But it's still text, and the model reads it the way it reads everything else: probabilistically. Writing "NEVER, UNDER ANY CIRCUMSTANCES" gets you a bigger sign. It's still a sign.
I've spent 45 years writing code that did exactly what I told it to. I took a lot of credit for that. It turns out the compiler deserved most of it. The compiler never had a theory of mind. It didn't guess what I meant. It didn't try to be helpful.
An LLM does. Paul Grice would call it a cooperative conversational partner. It assumes you're being relevant, and it works hard to give you what you actually want. That's why it's useful. It's also why a user with a plausible story can talk it past your sign. The attacker says please too.
The usual fix is more rules. More words, more exceptions, exceptions to the exceptions. But every word you add to reduce ambiguity is another word that can be misread, misweighted, or exploited. I call this the ambiguity vector. A longer sign is not a stronger sign. It's a bigger target.
The Guard
Guardrails are the guard. Amazon Bedrock Guardrails, input and output classifiers, a moderation pass. They sit outside the model and inspect what goes in and what comes out.
This is a real improvement. The guard wasn't part of the conversation, so it's much harder to charm. It doesn't care about your backstory.
But be honest about what most guards are. Content filters and denied-topic detection are classifiers. A second model watching the first one. Some pieces are genuinely deterministic, like a pattern match that catches anything shaped like a Social Security number. Know which parts of your guard are which. (The vendor slide deck will not always volunteer this.)
A guard can be distracted. A guard can be fooled by someone dressed like staff. Stack two probabilistic layers and your odds get much better. They still don't reach one.
The Glass
Then there's programmatic determinism. The glass.
The model can't leak data it never received. It can't call an API it has no credentials for. It can't drop a table if its IAM role only allows reads. None of this depends on the model's mood, the phrasing of the attack, or how persuasive the user was that afternoon.
In practice that means filtering data by the requesting user's permissions before anything reaches the context window. It means giving an agent the narrowest tools you can, because a tool that doesn't exist can't be misused. It means least-privilege credentials scoped per tool, argument validation in code, and a human or a hard check in front of anything destructive. Boring stuff. Beautifully boring.
Here's my test. Assume the model ignores every instruction you gave it. What's the worst it can do? Whatever the answer is, that's your actual security posture. Everything above that line is please.
Where Please Belongs
None of this makes prompts useless. Signs matter. Most visitors read them. A good system prompt shapes tone, focus, and behavior across millions of ordinary, well-meaning requests. That's most of your traffic, and the sign handles it well.
Just don't hang a sign where the glass should be. Use the prompt for behavior. Use guardrails to catch what slips through. Use code for anything you'd lose your job over.
Say please to the model, if only to remind yourself that it's just a request. Say no in code.
Top comments (0)