DEV Community

jsvnsk-tech
jsvnsk-tech

Posted on

Why AIs Seemingly Misbehave

Overview:

This article is about understanding why AI assistants and related AI systems seemingly misbehave. It explains why semantic understanding does not guarantee reliable instruction following. It examines why important constraints and factual requirements can be enforced more reliably. It offers guidance for setting realistic expectations when using AI systems.

keywords: AI assistants, large language models, LLMs, instruction-following, constraints, behavioral alignment, semantics, understanding, behavior, software engineering, invariants, agentic extraction pipelines, Retrieval-Augmented Generation, RAG, tool use

Understanding versus behavior:

Semantic capability does not guarantee reliable instruction following.

One of the strangest experiences with modern AI assistants is watching them fail at instructions that would be almost trivial for a human to follow. Consider a simple request: "Review this programming script, but leave all existing comments exactly as they are." The instruction is clear. There is little ambiguity. A human programmer would normally understand the boundary immediately: "Review or modify the code, but treat the existing comments as read-only." Yet an AI assistant may acknowledge the instruction, explicitly confirm that it understands it, and then proceed to modify the comments anyway. So why does this happen?

The problem is not necessarily that the AI failed to understand the words. Modern LLMs can process vocabulary, context, and many forms of natural-language meaning at a high level. Their ability to interpret language has improved substantially across many domains, although performance remains uneven, especially for ambiguous, specialized, or high-stakes tasks. The deeper problem is that understanding an instruction and reliably enforcing it are two very different things. A model can understand a request while still responding in a way that does not satisfy the user's intended objective.

This does not mean that LLMs cannot represent or follow constraints. Instruction following is a capability that modern models are specifically trained to perform. The issue is reliability: a model can follow a constraint successfully across many examples without providing the kind of deterministic guarantee that conventional software can provide.

The important distinction is therefore between capabilities that a model can perform probabilistically and guarantees that a surrounding system can enforce independently of the model.


Learned behavior versus explicit constraints:

An AI assistant's behavior is shaped by its learned parameters, post-training, the current context, and the instructions and constraints supplied by the surrounding system. It does not, by default, provide the same kind of explicit, externally verifiable constraint mechanism that conventional software can provide—for example, a representation in which some portions of a document are programmatically marked immutable and others editable.

When asked to improve code, a model may have learned strong patterns associating code improvement with improving documentation, formatting, and comments. Those learned tendencies can sometimes conflict with an explicit user constraint.

Consider the following interaction:

AI ASSISTANT: "I understand your explicit instructions. I will leave all of the comments unchanged."

HUMAN USER: "That's great. Please go ahead and fix the script."

AI ASSISTANT: Worked for 3.2 seconds

AI ASSISTANT: "OK, here is your updated program. I also changed or removed some of the comments."

The first statement above sounds like a commitment. Technically, however, that statement is itself generated by the model. It is evidence that the model represented the instruction in the current context, but it is not an independent guarantee that the subsequent output will satisfy the stated constraint.

A human might hear an instruction and establish a simple mental rule: comments → don't touch; code → may modify.

A language model can represent such a rule in context and may successfully follow it. However, the final artifact is generated through the model's inference process, with the instruction represented indirectly through its learned parameters and the current context. That process does not, by itself, provide a deterministic guarantee that a particular constraint will be preserved.

Unless the surrounding software provides an independently enforced restriction, there is no guarantee that the model will preserve a comment when generating the revised file.

The same distinction applies more broadly. A model may understand what the user wants while other objectives influence its behavior. Safety policies, system instructions, product constraints, and other optimization objectives can redirect a response. Platform-level policies can operate independently of stated user intent. Product and service design may introduce objectives and constraints beyond the immediate task, such as safety requirements, cost limits, rate limits, feature availability, or workflow requirements.


Institutional design:

Actual model behavior is shaped by several layers beyond the user's immediate request.

1. Training data + fine-tuning encodes learned patterns and preferences:

Training data contains a mixture of factual information, conventions, biases, and human judgments. Post-training processes can reinforce behaviors such as helpfulness, safety, refusal, and instruction following. These learned behaviors reflect choices made by model developers and cannot generally be overridden per conversation through user instructions alone.

2. Deployment constraints add enforced operational layers:

Rate limits can restrict access according to usage, capacity, or account-level policies. Feature gates can restrict capabilities based on account tier, product configuration, or availability. Authentication requirements can interrupt workflows when access to a particular service or capability is required. Architecture and infrastructure decisions are established outside any individual conversation.

3. Business logic sits underneath all surface interaction:

Product design can direct users toward particular workflows or capabilities. Commercial decisions determine which capabilities are available at different price points. Infrastructure costs can influence capacity limits, model selection, and usage restrictions.

Altogether, this is institutional and system design, not simply a semantic problem. Organizational decisions and product requirements can be incorporated into service architecture before a user begins a conversation. Commercial logic operates beneath the user-facing interaction. Safety and system-level constraints can impose boundaries that take precedence over a user's request.

None of these factors necessarily align with what the user wants in the moment. The issue is not simply whether the model understands the request. Rather, it is whether the system's actual behavior remains aligned with the user's objective after accounting for its training, policies, product configuration, and operational constraints.


Soft constraints versus hard constraints:

A useful engineering distinction is between soft constraints and hard constraints.

A prompt normally communicates a constraint to the model. The model may understand that constraint and attempt to follow it, but compliance remains part of the model's generated behavior.

A validator, schema, permission system, transaction boundary, or post-processing check can instead enforce a constraint independently of the model.

If preserving every existing comment is a hard requirement, the system should not merely ask the model to preserve comments. It should verify that the output satisfies that requirement.


The engineering approach to AI constraints:

Adding more and more instructions isn't the answer.

"Do not alter comments. Do not paraphrase comments. Do not correct comments. Do not reformat comments. Do not remove comments. Do not add comments. Preserve every character of every existing comment..."

Doing that shifts the burden from the machine to the human. That's backwards. The better engineering approach is to treat important requirements as invariants that can be verified, rather than promises the model is merely asked to remember.

For example, a coding system could identify the original comments, allow the AI to modify the executable code, and then automatically compare the resulting comments against the originals. If any comment changed, the system could reject or revert the modification. The user wouldn't need to cosplay as a lawyer and write an airtight contract. The software would enforce the obvious boundary.

This principle extends beyond code editing. When correctness depends on exact data, the workflow should move retrieval and validation into systems whose behavior and outputs can be independently tested and verified. The LLM can then be used for tasks such as context-constrained parsing, transformation, or interpretation, while programmatic checks enforce requirements that must not be violated. This is especially important when exact preservation, formatting, data integrity, or other hard constraints matter.


Factual data-based tasks:

An AI assistant is not a database.

It is equally important to distinguish language generation from authoritative data retrieval. LLMs are probabilistic models that generate text based on learned patterns and the context provided at inference time. This differs from conventional database systems, which can provide structured retrieval semantics and, depending on the system, transactional or consistency guarantees.

Although LLMs can recall a large amount of factual information, their internal knowledge is not a structured, authoritative, or consistently verifiable database. They can produce incorrect or fabricated details, particularly when asked for precise dates, numbers, quotations, citations, or other granular facts. Users should therefore not treat an AI assistant as a flawless source of structured factual data when authoritative retrieval is available and correctness matters.

A database should not be confused with an infallible source of truth either: databases can contain incorrect, incomplete, or stale data. The important distinction is that a database provides an explicit retrieval mechanism whose behavior and results can generally be tested independently of the language-generation process.


Agentic extraction pipelines:

To automate fact-based tasks more reliably, it is often better to build a pipeline in which programmatic systems retrieve the source data, and the LLM is used for context-constrained parsing or transformation. Retrieval-Augmented Generation (RAG) is one related architecture in which external information is retrieved and supplied to the model as context. "Tool use" is a broader concept in which a model invokes external systems to retrieve information or perform actions. Neither approach makes the LLM inherently deterministic, but both can constrain the information available to the model and move critical retrieval or verification steps into systems that can be tested directly.

Retrieval also does not guarantee a correct answer. The retrieved source may be incomplete, outdated, or incorrect, and the model may still misinterpret or misapply the retrieved information. The advantage is that retrieval and validation can be moved into independently testable components rather than relying entirely on the model's parametric memory.

For fact extraction, a robust pipeline can include the following components:

  • deterministic or otherwise independently testable retrieval
  • context-constrained parsing
  • programmatic verification layer
  • flag for human intervention, if merited

An example of an extraction pipeline to obtain the U.S. release dates of individual song titles from an API:

  • Use a script, such as Python with an API, or SPARQL queries, to retrieve the relevant source data for the specific song page.
  • Pass the retrieved text to the LLM using a strict prompt; for example: "Extract the U.S. single release date from the provided text. Format as YYYY-MM-DD. If the text does not explicitly state a U.S. release date, output 'NULL'."
  • Apply a programmatic check to verify that the extracted value conforms to the expected format and can be traced to the relevant source passage.
  • Where the distinction requires semantic interpretation—for example, distinguishing a U.S. single release from an album release—the result may still require model-based reasoning or human review.
  • If the date does not match the source data, or the source does not explicitly establish a U.S. release date, flag the result for human review.

Fluency does not equate with compliance:

This illustrates a broader limitation of current AI assistants: fluency can create the appearance of reliable instruction-following without actually providing it. An assistant can sound completely certain that it understands you. It can explain your requirement back to you perfectly. It can even apologize after violating it. None of those things guarantee compliance. The same limitation appears when an assistant is asked to provide precise factual data without access to, or verification against, an authoritative source.

For many tasks, these distinctions don't matter much. But when exact preservation, formatting, data integrity, or factual accuracy matter, they matter enormously. The issue isn't that AI assistants are incapable of understanding simple instructions. Instead, the key point is that conversational understanding should not be confused with dependable execution. Similarly, the ability to produce plausible factual information should not be confused with authoritative data retrieval.

A good assistant shouldn't require the user to become increasingly precise merely to prevent it from "helpfully" changing something that was explicitly declared off-limits. If a requirement is important enough to be treated as a hard boundary, the system should represent and enforce that boundary explicitly rather than relying solely on conversational instructions. When a requirement is important, however, the safest approach is to encode it in the surrounding system and verify the result. Deterministic retrieval where appropriate, constrained model use, programmatic validation, and human review can provide stronger guarantees than conversational instructions alone.

The practical point is therefore not to abandon AI assistants, but to use them according to their actual strengths and limitations. Use language models for interpretation, transformation, and other tasks where probabilistic generation is appropriate. Use external tools and authoritative sources when exact information is required. Use invariants and automated checks when particular parts of an artifact must remain unchanged. And, where correctness matters, verify that the machine actually did what it was told.

Top comments (0)