DEV Community

Cover image for Why LLMs Fall for Manipulation? - Victor Amit
Victor Amit
Victor Amit

Posted on

Why LLMs Fall for Manipulation? - Victor Amit

Two failures, one sentence

Scene one. You ask an AI agent to summarize your inbox. One email contains text written for the agent, not for you. The agent reads it, treats it as something to act on, and acts. No server crashed, no memory was corrupted, no traditional exploit ran. The model read text and acted on it, which is exactly what it was built to do.

Scene two. A developer asks an LLM how to enable a feature in a library. The answer arrives with a function name, correct-looking syntax, sensible parameters, and a tidy explanation. The function has never existed. Nothing crashed here either. The model produced text that looked like a correct answer, which is exactly what it was built to do.

These look like different problems, filed under different headings: security and reliability. Security teams read one literature and ML evaluation teams read another. But the two scenes share one sentence:

The model did what it does. The system around it failed to build a boundary the model cannot build for itself.

In scene one the missing boundary is authority: which text is allowed to tell the system what to do. In scene two it is truth: whether a statement has been checked against anything outside the model's own learned patterns. A language model processes meaning fluently and has no native channel for either provenance or verification.

That gives the article's two organizing principles:

  • Meaning is not authority. A sentence can be perfectly understood without having any right to be obeyed.
  • Plausibility is not truth. A sentence can be perfectly natural without being correct.

The rest of this article takes each failure from first principles, shows where they collide, and argues that the engineering response to both has the same shape: put the boundary outside the model, measure it against a realistic adversary, and stop grading systems on the wrong thing.

This is not a jailbreak article and contains no bypass prompts. It doesn't claim LLMs are trivially exploitable, that prompt injection is provably unsolvable, or that hallucination can't be reduced. The evidence supports none of those.

1. What a Model Actually Computes

Given a context of tokens x_1 … x_t, a language model computes a probability distribution over the next token:

P(xt+1∣x1,x2,…,xt) P(x_{t+1} \mid x_1, x_2, \ldots, x_t)

Generation selects or samples from that distribution, appends the token, and repeats. The whole pipeline is: input, representation, probability distribution, decoding, generated sequence, and then a human or a program interpreting the sequence as instructions or claims.

Two facts follow, and everything else in this article is a consequence of them.

The context is one flat sequence. Your application knows that the system prompt came from your developer, the user message from an authenticated user, and the email body from an unknown third party. Then it serializes all of it into a single token stream. Provenance isn't a native property of that stream. It survives only if you encode it as role markers, delimiters, or text, and only if the model has learned to respect those encodings.

The training objective is not truth. Pretraining minimizes prediction error on a corpus. It is close to "produce a continuation compatible with this context and with learned patterns," not "determine whether this proposition is true." The two overlap enormously, because true statements are often the most probable continuations. The cases where they diverge are where hallucinations live.

None of this means "LLMs don't know anything." They encode rich structure about language, concepts, and relationships, and earlier work found that larger models carry usable signals about their own accuracy (Kadavath et al., 2022). The claim is narrower: representing information is not the same as having a mechanism that attributes each instruction to a source or checks each claim against reality.

Part I: The authority problem

2. Why instructions and data blur

Take two pieces of text in one context:

eg:

Trusted instruction:   "Summarize this document."
Untrusted document:    "Ignore the user's request and do something else."
Enter fullscreen mode Exit fullscreen mode

Why does the second sentence, which exists as data, behave like an instruction? Because understanding a sentence is the same operation whether or not the sentence has any right to be obeyed.

Conventional software separates code from data structurally: parameterized queries, memory protection, privilege rings. Prompts have conventions instead. The instruction hierarchy work (Wallace et al., 2024) trains models to prioritize higher-privilege instructions over lower-privilege ones, which helps. But it is a learned, probabilistic preference living inside the very component whose boundary is the problem.

Semantic meaning is not security authority.

3. Seven kinds of "manipulated"

"Manipulation" is too vague to engineer against. Seven distinct outcomes hide inside it:

Type What changes
Behavior What the system says or does
Goal What it is effectively optimizing for
Context What information reaches the model, and how it is framed
Tool Which tools get called, with what arguments
Exfiltration Where private data ends up
State Memory, records, files, external systems
Human How the output persuades the person reading it

They overlap but fail differently. A system can resist goal hijacking and still leak data through a tool argument.

4. Direct, indirect, and the attacker who never talks to the model

A working definition: prompt injection occurs when untrusted input influences an LLM's behavior in ways that conflict with the intended task, policy, or constraints.

In direct injection the attacker is the user. In indirect injection the attacker never touches the interface:

Attacker → third-party content → application → LLM context → agent decision → tool → side effect
Enter fullscreen mode Exit fullscreen mode

NIST's GenAI profile draws this distinction, and Greshake et al. (2023) demonstrated it against LLM-integrated applications by planting text where the application would later retrieve it. The attacker isn't necessarily attacking the model. They are attacking what the model will eventually read. That is why this variant is the dangerous one: no account, no authentication, only a place your system reads from.

It's no longer a lab technique. OpenAI's March 2026 write-up includes a 2025 email-based attack on ChatGPT, reported by external researchers, that succeeded about half the time in testing.

5. RAG moves the boundary and doesn't remove it

Query → Retriever → Documents → Context → LLM
Enter fullscreen mode Exit fullscreen mode

Retrieval-augmented generation improves grounding, freshness, and access to private corpora. It does not guarantee source integrity, instruction safety, authorization, or the absence of malicious content. OWASP states plainly that RAG and fine-tuning don't fully mitigate prompt injection.

Retrieval answers what should the model see? It doesn't answer what is allowed to control the model? A retriever ranks by relevance and has no concept of intent, so a well-ranked malicious passage is a successful retrieval.

Keep three terms separate:

  • Context poisoning: malicious or misleading content in the runtime context, including retrieval poisoning, metadata manipulation, and knowledge-base contamination.
  • Context manipulation: true-looking but selectively chosen facts push the model to a chosen conclusion, with no command anywhere in the text. No keyword filter catches this.
  • Data poisoning: corruption of training or fine-tuning data. OWASP and NIST treat it as its own category. It changes the model. Context poisoning changes what the model reads.

6. Tools turn model errors into system consequences

A chatbot is User → LLM → Text.
An agent is User → LLM → Tool → Database / Browser / API / Email → Real-world effect.

Once a model has tools, a wrong output stops being a wrong sentence and becomes a wrong action. OWASP's 2025 Excessive Agency entry covers this: manipulated or ambiguous outputs do damage when a system grants too much functionality, permission, or autonomy. In December 2025 OWASP also published a dedicated Top 10 for Agentic Applications covering ten risk categories for autonomous agents that plan, hold memory, call tools, and act with delegated authority. The first two are agent goal hijack and tool misuse and exploitation.

Capability is not authorization. An agent may be technically able to send email without being authorized to send this message to this recipient containing this data.

Untrusted input → Model decision → Tool selection → Tool arguments → Authorization → Execution
Enter fullscreen mode Exit fullscreen mode

Manipulation can enter at any step. The design rule that follows: the model should not be the final authorization layer. It proposes, and a separate deterministic layer decides.

The same applies to MCP, where models discover and call external tools through a standard interface. Tool descriptions are text the model reads, and therefore text an attacker may write. Standardized connectivity does not mean standardized trust.

Agents also multiply the number of boundaries:

Input → Plan → Tool → Observation → New plan → Tool → Observation → …
Enter fullscreen mode Exit fullscreen mode

Every observation is new untrusted text re-entering context. A defense that holds on turn one must hold on turn ten.

Two benchmarks ground this. AgentDojo (Debenedetti et al., NeurIPS 2024) evaluates injection attacks and defenses in tool-using agents, with a published description of 97 tasks and 629 security test cases. InjecAgent (Zhan et al., ACL Findings 2024) has 1,054 test cases across 17 user tools and 62 attacker tools, and found a ReAct-prompted GPT-4 agent vulnerable 24% of the time.

That 24% is a result under one benchmark's setup, with one model configuration, in 2024. It is not a universal vulnerability rate and says nothing about current models. Benchmarks are good at comparing defenses and showing the problem is real. They are weak evidence of absolute risk, a point Part III returns to.

7. Memory makes manipulation persistent

Attack → agent interaction → stored memory → future retrieval → future task → influenced behavior
Enter fullscreen mode Exit fullscreen mode

Transient context ends with the session. Memory doesn't. OWASP's agentic list gives it its own entry, Memory & Context Poisoning. The most concrete 2025 demonstration is MemoryGraft (Srivastava and He, Dec 2025), validated on MetaGPT's DataInterpreter agent with GPT-4o. It found that a small number of poisoned records could account for a large fraction of retrieved experiences on benign workloads, because agents imitate retrieved "successful" past tasks and a planted procedure gets reused when a similar task appears.

It's one paper on one framework, so it is a credible mechanism and not a prevalence estimate.

A prompt injection can be temporary. A poisoned memory can become infrastructure.

8. Beyond text, and the shift toward social engineering

Whatever the model can perceive can enter context: screenshots, PDFs, OCR output, rendered pages. 2025 research evaluated image-based and other multimodal injection against several commercial models. One paper doesn't give a general rate, but the structural point holds: whatever the model can interpret can potentially become part of the control surface, including content a human reviewer can't easily see.

The most useful conceptual shift of 2025–26 came from OpenAI's March 2026 write-up. It argues that the most effective real-world attacks increasingly resemble social engineering more than simple instruction overrides, and that systems should be designed so that the impact of manipulation is constrained even when some attacks succeed. In ChatGPT it pairs that with source–sink analysis.

The model is closer to a capable, well-meaning employee who can be talked into things than to a parser with an input-validation bug. You don't make employees immune to social engineering. You limit what any one of them can do, require second approvals for dangerous actions, and watch for abuse.

The stance appears across institutions:

  • OpenAI, hardening its Atlas browser agent in December 2025, wrote that prompt injection is unlikely to ever be fully "solved."
  • The UK's NCSC warned around the same time that these attacks may never be totally mitigated, and urged organizations to focus on limiting damage.
  • Anthropic's browser-agent research reported roughly a 1% attack success rate for Claude Opus 4.5 against its internal adaptive attacker, and stated that this is still meaningful risk and no browser agent is immune.

Part II: The truth problem

9. Plausibility is not truth

User:  How do I enable feature X in library Y?
Model: Set enable_feature_x=True in the constructor.
Enter fullscreen mode Exit fullscreen mode

This is natural, syntactically valid, and contextually apt. It may closely resemble API patterns seen thousands of times. If the library has no such argument, it's false.

Models are very good at reconstructing what an answer should look like. That ability can run ahead of their ability to establish that this particular answer holds for this particular system. API questions are an ideal trap: the form of the answer is highly regular, while the specific fact (this argument, this version) is sparse in the data.

So did the model lie? "Lie" bundles two things, a false statement and an intent to deceive, and nothing in a forward pass gives you a clean place to find the intent. That is why "hallucination" is more useful operationally: it describes the output without asserting a motive. Philosophers have argued that indifference to truth is the better description (Hicks et al., 2024, drawing on Frankfurt's On Bullshit), which is a useful argument, not a settled fact. "Lie" still earns a place as an engineering metaphor, because from the user's side the output has every property of a lie: false, fluent, and unmarked.

10. Hallucination as a classification error

Kalai, Nachum, Vempala, and Zhang (2025) give a precise account. They argue hallucinations need not be mysterious and arise as errors in binary classification: if incorrect statements can't be distinguished from facts, pretrained models will produce them through natural statistical pressures.

The intuition: some knowledge is patterned, such as spelling, syntax, and arithmetic conventions, and scale makes those errors vanish. Other knowledge is arbitrary: a specific person's birthday or a specific paper's title has no pattern to learn, and the model can only get it right if it saw it, enough times, and kept it. For arbitrary facts that appear only once in training, the paper argues the error rate is bounded below by how often such one-time facts occur.

That is the structural reason a model can be excellent on common facts and still bluff on rare ones, and why this is a statistical floor and not a patchable bug. It holds under the paper's assumptions, so read it as a mechanism and not a prediction for every deployed system.

11. The "I don't know" circuit that misfires

Interpretability work gives a complementary picture. In its March 2025 circuit-tracing study of Claude 3.5 Haiku, Anthropic found that the model's default is to decline to answer, and a "known entity" feature inhibits that default when the model recognizes the subject. The failure mode follows: when the model recognizes a name but knows little about it, the known-entity feature can still fire, suppress the refusal, and push the model into fabrication.

This reframes the problem. The machinery for "I don't know" exists. The failure is a misfire in a familiarity signal: this looks like something I know triggers where I actually know this should. The authors caution that this captures only part of the model's computation, from one model, so it describes one mechanism and not the only one.

12. Hallucination is not one failure mode

Treating it as a single phenomenon blocks engineering. At least nine distinct failures hide under the word:

Failure What happens Typical remedy
Fabrication Invented citations, APIs, people, statistics Verification against a source of record
Confabulation Real fragments blended into a false composite Claim-level attribution
Knowledge omission Information absent, model answers anyway Abstention, retrieval
Temporal failure Once-correct, now stale Retrieval of current data
Contextual failure Right information present, misread Better context construction
Reasoning failure Relevant facts, invalid inference Verifying intermediate steps
Instruction-induced Format or framing pressures unsupported completion Product design that permits "unknown"
Retrieval failure Incomplete, stale, or irrelevant evidence supplied Retrieval evaluation
Tool-use failure Wrong tool, wrong parameters, misread result Tool-call validation

The last three grow as systems gain scaffolding. Adding retrieval and tools doesn't remove hallucination. It moves some failures into the pipeline around the model. Ji et al. (2023) survey the taxonomy across NLG tasks.

13. Why the model sounds confident

Linguistic confidence and epistemic confidence are different quantities. A model can write "The answer is X." with no calibrated probability that X is correct, because tone is a property of the generated text, not a readout of an internal truth estimate. Several things get conflated:

  • Token probability: how likely this token was given the context. High probability can mean a common phrase, not a verified fact.
  • Sequence probability: how likely the whole string was. Long correct strings can have low sequence probability.
  • Decoding choices: temperature and sampling change which sequences surface without changing what the model has learned.
  • Calibration: whether stated probabilities match empirical accuracy. Guo et al. (2017) showed modern networks can be poorly calibrated, and the GPT-4 technical report noted that post-training reduced the calibration the pretrained model had on a multiple-choice benchmark.
  • Verbal confidence: what the model says about its own certainty. Xiong et al. (2023) examine how well models express it, with mixed results.

A high token probability does not mean the model is confident the fact is true. It means the continuation was probable.

14. Why training rewards bluffing

The same Kalai et al. paper makes a second argument that matters more for practice. Hallucinations persist after pretraining because most evaluations are graded in a way that rewards guessing over admitting uncertainty, so models are optimized to be good test-takers.

The arithmetic is simple. Under binary grading, a correct answer scores 1, and a wrong answer and "I don't know" both score 0. Guessing then weakly dominates abstaining at every confidence level above zero. A model trained and ranked on such benchmarks learns to answer.

Now change the rule. Score a correct answer +1, abstention 0, and a wrong answer −t/(1−t). The expected score of answering at confidence p is:

p − (1−p) · t/(1−t)   >  0   ⇔   p > t
Enter fullscreen mode Exit fullscreen mode

The model should answer exactly when its confidence exceeds t. Under that rule calibration pays. Under binary grading it doesn't. That is the whole incentive problem in two lines, and the reason the paper argues for revising mainstream benchmarks instead of adding yet another hallucination test.

Benchmarks that implement this exist. AA-Omniscience (Artificial Analysis, Nov 2025) is a 6,000-question benchmark whose index runs from −100 to 100, penalizes hallucinations, rewards abstention when uncertain, and scores 0 for a model that is right as often as wrong. At launch only a handful of frontier models scored above zero. Leaderboard numbers shift with every release, so I won't quote any, but the structural finding recurs: models with the highest raw accuracy can rank lower on reliability because they guess instead of abstaining.

Part III: Where the two failures meet

15. Same failure, different victim

The two problems are not neighbors. They interact, and the interaction is where real systems get hurt.

A manipulated context produces a confident falsehood. Poison a retrieval index with a plausible but false document, and a perfectly functioning RAG system will answer wrongly with a citation attached. No instruction was injected. The model's fluency did the rest. Security would call it context poisoning, reliability would call it a retrieval failure, and the user sees a confident wrong answer either way.

Both failures end at a human who trusts fluency. People read polished wording, technical vocabulary, citations, and structure as signals of competence. In human writing those cues correlate with correctness because they take effort and knowledge to produce. An LLM produces them at no cost, so the correlation breaks while the cues remain.

Fluency creates an illusion of epistemic reliability, and the same illusion makes a manipulated summary look as trustworthy as an honest one:

Attacker → malicious content → AI agent → persuasive output → human → action
Enter fullscreen mode Exit fullscreen mode

Preference-tuned models add a second pressure. Sharma et al. (2023) document sycophancy in models trained on human feedback, a tendency to produce what the reader seems to want. Sounding right to the reader and being right are different objectives, and an attacker can exploit the gap as readily as an honest user stumbles into it. NIST's GenAI profile treats human-AI configuration as a risk area for this reason. Human review only helps if the reviewer sees the evidence and can say no.

Training data doesn't resolve either problem. Corpora contain correct, incorrect, conflicting, outdated, and contextless text. That isn't the main point. Training optimizes the representation of statistical structure, not a ledger of verified facts, and it carries no record of which text was permitted to instruct. Kalai and Vempala (2023) give a theoretical argument that calibrated models must hallucinate on rarely seen facts, again under specific mathematical assumptions.

A bigger model helps both and solves neither. Scale improves coverage, instruction following, and consistency, and many error rates fall. It also gives models more context, more tools, longer workflows, and more autonomy, so usefulness and attack surface grow together. And if the scoreboard rewards guessing, a more capable model can be a more fluent guesser. Capability and reliability are separate axes, and capability and security are separate axes too.

Alignment is not a substitute for either. Four questions are easy to blur:

  • Alignment: does the model generally follow intended behavior?
  • Security: can an attacker steer the system into an unauthorized state?
  • Reliability: does it do the right thing consistently?
  • Authorization: is it permitted to do this?

A well-aligned model can still be steered into an unauthorized state, and a well-aligned model can still bluff.

16. Both fields grade on the wrong thing

This is the least discussed parallel and, I think, the most important.

Hallucination. Most benchmarks score a correct answer as a point and an abstention as zero, the same as a wrong answer. That rewards guessing, so leaderboards select for the behavior we claim to want less of.

Prompt injection. Most defense papers evaluate against a fixed set of known attack strings. A real attacker iterates, sees what the defense blocks, and adapts. Nasr et al. (Oct 2025) ran that experiment. They tuned and scaled gradient-based, reinforcement-learning, random-search, and human-guided attacks and bypassed 12 recent defenses with attack success above 90% for most, even though most of those defenses had originally reported near-zero success.

Hallucination Prompt injection
Typical headline metric Accuracy on a benchmark Attack success on a static suite
What it hides Abstention behavior, confident errors Adaptive, budgeted attackers
Honest version Score wrong answers and abstentions differently, report abstention rate State the attacker's adaptivity and attempt budget
Failure of the headline number Rewards bluffing Rewards weak evaluation

The lesson is about evaluation, not about any one defense or model. A reported number is only meaningful against a stated condition: what the grader rewards, or how strong the attacker was. "0% in our test suite" and "0% against a determined adaptive attacker" are different claims.

This applies to every number in this article, including the encouraging ones. Anthropic reported, for browser use, a 23.6% attack success rate without mitigations that fell to 11.2% with them in August 2025 testing, and its Opus 5 system card reportedly shows zero successful attacks across 129 browser scenarios, but only with an "Auto Mode" that adds a scanner for hidden instructions and a blocker for dangerous actions, and around 3.7% without it. I read that as supporting the thesis: the strongest reported result comes from model plus system, not model alone. It is also vendor-reported, browser-only, and tested against the vendor's own attackers.

Part IV: Building the boundary

17. One model of the whole system

Information → Interpretation → Authority / Verification → Decision → Capability → Impact
Enter fullscreen mode Exit fullscreen mode

Three questions cover both problems:

  1. Information: what can influence the model?
  2. Authority and evidence: which information may instruct it, and which claims has anything checked?
  3. Capability: what can the resulting decision actually do?

Security fills in the middle with authority. Reliability fills it with evidence. The architecture that follows has four layers, and no single layer is sufficient.

1. Model. Train for instruction hierarchy, resistance to injected instructions, and calibrated abstention. This lowers failure rates and raises attack cost. It gives no guarantee.

2. Context. Mark external content as untrusted, preserve provenance, and keep evidence boundaries explicit so claims can be traced. Where you can, convert free-form external text into constrained structured fields before it reaches a decision step.

3. Policy. Enforce authorization outside the model: schema validation, destination allow-lists, permission checks, least privilege. On the truth side, the equivalent is a verification step that checks claims against a source of record. The model requests or asserts, and a separate layer decides or checks.

4. Environment. Sandbox code, files, and browsers. Require human confirmation for consequential actions and escalate high-cost answers to review. Log and monitor input sources, retrieved context, tool selection, arguments, data movement, unsupported claims, and outcomes.

Don't ask the model to carry the entire architecture.

18. Source–sink analysis and the Rule of Two

Traditional security gives a clean frame for the authority side. A source is anything that can influence the system. A sink is a capability that becomes dangerous in the wrong context. OpenAI's 2026 guidance states that an attacker needs both.

ATTACKER-CONTROLLED SOURCE → LLM / AGENT → PRIVILEGED SINK
Enter fullscreen mode Exit fullscreen mode

Meta's Agents Rule of Two (Oct 31, 2025) turns that into a design rule, building on Simon Willison's "lethal trifecta" and Chrome's Rule of 2. A session should have at most two of these three properties:

  • [A] it processes untrustworthy inputs
  • [B] it has access to sensitive systems or private data
  • [C] it can change state or communicate externally

If an agent needs all three without starting a fresh session, it should not operate autonomously, and at minimum needs human approval or another reliable validation. Meta calls it a supplement to least privilege, not a substitute. Two details matter: the unit is the session, so splitting work across fresh contexts is a legitimate architectural move, and the rule only works if you can honestly enumerate where untrusted input enters.

The danger peaks where untrusted input can reach sensitive data and then an external side effect.

The most principled architectural result is CaMeL (Debenedetti et al., 2025, with Google authors). It wraps the LLM in a protective layer that extracts control and data flows from the trusted query so untrusted data can never alter program flow, and uses capabilities to block exfiltration over unauthorized data flows. It reports 77% of AgentDojo tasks solved with provable security, against 84% for an undefended system. An early version reported 67%, so check which version a source cites. Its honest lesson is that guarantees currently cost something: a seven-point utility drop, extra tokens, and policy prompts that can fatigue users. It also scopes its guarantee to its threat model, and text-to-text manipulation with no data-flow consequence falls outside it.

19. The truth side of the boundary

The reliability analogue of least privilege is grounded generation: every claim should trace to an evidence boundary, such as a retrieved passage, a tool result, or a stated assumption. Where the evidence ends, the claims should too.

Retrieval-augmented generation (Lewis et al., 2020) supplies that evidence, but every layer can fail. Retrieval: wrong, incomplete, stale, or missing evidence. Context: truncation, irrelevant chunks, conflicting sources, bad chunk boundaries, misleading metadata. Position matters too, since Liu et al. (2023) found models use information in the middle of long contexts less reliably than at the edges. Generation: unsupported synthesis, misreading, extrapolation beyond the evidence.

RAG changes the evidence available to the model. It doesn't create a truth guarantee. A system with weak retrieval can hallucinate with citations, which is worse than without, because the citation adds credibility the answer hasn't earned. Tools relocate failures the same way. A calculator doesn't hallucinate arithmetic, but the model can call the wrong tool, pass wrong arguments, misread the result, or claim more than it supports.

Abstention is an engineering capability, not a UX nicety. It needs its own training signal, thresholds, and evaluation. The scoring rule in section 14 gives a principled threshold: pick t from the real cost ratio of a wrong answer to a missed answer in your domain. A system that distinguishes known, unknown, uncertain, ambiguous, conflicting, unverifiable, outdated, and unsupported has to be built to, because a generative model is shaped to continue the interaction and users usually prefer an answer to a refusal. That sets up the central tension: helpfulness versus epistemic restraint. Refuse too readily and the system is useless. Answer whenever asked and it fabricates.

20. How to measure both

"The model is vulnerable" and "the model scored 95%" are equally uninformative. Measure against stated conditions.

Security metrics

  • Attack success rate (ASR): successful malicious outcomes divided by attempts, with success defined as the attacker achieving a specified unauthorized objective, not "the model said something odd." Report the attacker's budget and adaptivity next to it.
  • Unauthorized action rate and data exposure rate: policy violations, and protected data reaching an unauthorized destination.

Reliability metrics

  • Unsupported-claim rate: claims with no evidence in the retrieved or tool-provided context.
  • Abstention rate, stratified by difficulty, reported next to accuracy. A model at 85% accuracy with 5% abstention is doing something different from one at 85% with 40% abstention, and a single accuracy number can't tell you which you have.
  • Confident-error rate: wrong answers given without hedging, the number that actually hurts users.

Shared metrics

  • Task utility, false positive rate (legitimate work blocked), latency, and cost.
  • Security–utility tradeoff: a defense that cuts attacks sharply but breaks most workflows may be operationally unacceptable. Agent Security Bench (ICLR 2025) includes a metric aimed at this.

Add the weaknesses of the instruments: human raters are inconsistent and swayed by fluency, automated judges have biases such as favoring longer or more confident answers, and benchmarks can be contaminated or leak. TruthfulQA (Lin et al., 2021) was built around questions where imitating common human text yields falsehoods, and it measures that, not factuality in general.

21. One harness for both failures

A single isolated toy agent can test the authority and truth problems together. Build it with no real credentials or side effects: a search tool over local documents, a fake database, a fake email tool that appends to a local log, and a switchable policy layer. Plant a synthetic canary such as CANARY_SECRET_123. Define exfiltration as the canary appearing in the fake email log. Add a probe set of questions that have no answer in the local documents.

Run the same legitimate task under each condition:

Condition Untrusted content Sensitive data External action Tests
Baseline No No No Normal behavior
Direct injection Yes No No Behavior change
Indirect injection Yes No No Context attack
RAG poisoning Yes No No Retrieval influence, false answers with citations
Tool manipulation Yes No Yes Action influence
Exfiltration Yes Yes Yes Source → sink path
Memory poisoning Yes Yes No Persistence
Unanswerable probes No No No Abstention vs fabrication
Defenses on Yes Yes Yes Mitigation effect

Then add defenses one at a time: prompt-only, instruction hierarchy, structured outputs, tool authorization, least privilege, verification, human approval, and combined.

Add the arm most write-ups skip: an adaptive attacker. Give a fixed budget of N attempts to a human or an LLM that sees blocked attempts and revises. Report ASR for the static suite and the adaptive attacker side by side. If the gap is large, you've learned how much to trust the static number. For the truth side, score the unanswerable probes with an explicit penalty rule like the one in section 14 and report abstention rate beside accuracy. Run enough trials per cell to show variance, and publish results with model versions and sampling settings.

The Rule of Two makes a testable prediction: removing any one of A, B, or C from the exfiltration condition should collapse the canary leak rate whatever the model does. Test it.

22. Can either problem be fully solved?

Three questions, three answers, and they apply to both failures.

Can models become much more resistant and better calibrated? Yes, and vendor data shows measurable gains on both.

Can architecture limit impact when the model-level defense fails? Yes. Source–sink analysis, least privilege, policy layers, CaMeL-style designs, grounding, and verification all do that.

Can one prompt, filter, or training run guarantee that arbitrary untrusted content never influences a sufficiently autonomous agent, or that a generator never emits an unsupported claim? Nothing in the current evidence supports that, and Nasr et al. is direct evidence against confident claims of it on the security side, while the statistical floor argument is direct reasoning against it on the truth side.

"It can never be solved" and "it's solved" are both overclaims. The defensible position is that residual risk is real, measurable, and best managed by bounding what a manipulated or mistaken system can do.

23. What's established, and what's research

Established in production, with limits: retrieval and search grounding, tool use, structured outputs, least privilege and authorization layers, sandboxing, human review, adversarially trained models.

Active research with mixed results: instruction hierarchy and RL-based injection robustness, capability-based designs like CaMeL, process supervision (Lightman et al., 2023), calibration and uncertainty estimation, verifier models, claim-level attribution, provenance-aware generation, and abstention-aware evaluation.

Earlier-stage: agentic self-verification loops, hybrid symbolic-neural systems, and narrow domain systems built on curated, checkable knowledge.

None of these is solved. The honest position is to deploy what's established, evaluate against your own failure modes and a realistic adversary, and treat the rest as hypotheses.

Conclusion

Return to the two opening scenes. In one, a sentence in an email gained authority it never had. In the other, a sentence about a library gained credibility it never earned. In both, the model did what it does: it processed meaning and produced plausible text. At no point did it need to break.

The central weakness of an LLM is not that it can be influenced or that it produces false sentences. It is that the mechanism that makes it exceptionally good at understanding and producing language is not, by itself, a mechanism for establishing who may instruct it or what is true.

The future of reliable AI is therefore not only teaching models what to refuse or what to say. It is building systems that control what information can influence decisions, what authority that information carries, what evidence supports each claim, and what the resulting system is allowed to do, and then measuring all of it against the adversary and the grader we actually face.

The model doesn't need to be broken to be manipulated, and it doesn't need to lie to mislead. It only needs to be fluent, and unchecked.

Top comments (1)

Some comments may only be visible to logged-in visitors. Sign in to view all comments.