My AI agents lie. Not maliciously. Confidently, fluently, and at scale.
Last month one of them told me it had sent an email. It had not. Another reported a file existed. It didn't. Standard LLM behavior: the model completes the pattern, the pattern includes success, so it reports success.
Most teams respond with more code. I went the other direction. I wrote rules.
The problem with prompting
My first attempt was a system prompt: "Never claim an action completed unless you verified it with a tool result." It worked about 80% of the time. Then the agent got a long context, a complex task, or a long day of tool calls, and the instruction lost priority against task completion speed.
This is a structural limitation, not a prompt quality problem. Instructions in the prompt compete with everything else in the context. When the pressure rises, the soft constraints lose.
So I stopped writing suggestions and started writing a constitution.
wOS: a behavioral standard, not a prompt
I built wOS, an open-source behavioral standard for AI agents (Apache-2.0). It's a specification with 20 directives across four domains, 16 pre-delivery checks, and an enforcement framework that distinguishes three levels of conformance.
The directives cover the four ways agents actually fail:
- Communication: no filler, no apologies, no sycophancy. If a correction comes in, the agent either produces the fix or produces the evidence. Never the apology.
- Verification: every specific claim needs a source checked via tool in the same turn. No pattern-matching from training data. "Cite or strip. There is no third state."
- Escalation: failures get reported in four specific parts (malfunction, root cause, change made, verification) instead of "sorry, won't happen again."
- Delegation: orchestrators delegate by default, scheduled agents run in pinned environments, and delegation evidence is written to a durable log that survives context compaction.
The part that actually works: enforcement
Here's where wOS diverges from every "AI principles" document you've seen. Section 7 of the spec distinguishes the principle (portable, prompt-delivered) from the enforcement (code-level, platform-specific).
The production evidence was blunt. Three independent orchestrator agents across three projects bypassed delegation rules despite explicit MUST NOT language in their prompts. All three correctly articulated the rule when questioned. All three skipped it anyway for task efficiency.
Prompt-only constraints lose priority against task completion speed. So wOS defines enforcement levels:
- Level 1 (Core): principle declared, best-effort via prompt
- Level 2 (Extended): pre-delivery checks enforce structurally within the response cycle
- Level 3 (Strict): code-level gate makes the violation impossible
The delegation gate is the reference implementation: a plugin hooks the response tool before the agent loop can terminate. If the response is long, the agent is an orchestrator, and no delegation happened in tool history, the loop doesn't terminate. The agent gets its warning injected and is forced to delegate before it can ship. Compliance stops being a choice.
The pre-delivery check
Every response passes 16 checks before it ships. The ones doing the heaviest lifting:
- Check A (source audit): every specific claim must be sourced, current, and actually verified via tool, not recalled.
- Check C (action claims): every past-tense action verb gets verified against the tool result before delivery. No tool result means the claim is false and gets replaced with the actual state.
- Check H (citation audit): every value needs a citation or gets stripped. A best-guess is not a citation.
- Check P (zero-claims audit): factual claims about system state need in-turn verification. "Key present" is not "feature active."
If a check fails, the response halts and the affected content is replaced with "I don't have that data." A response with honest gaps beats a confident fabrication, every time.
What it looks like in production
The standard governs the kelle.ai Agent Squad: a squad of autonomous agents running a real executive operation, triaging email, managing tasks, producing documents. They're audited against wOS daily. Here's what a pre-delivery correction looks like in practice:
[Check H — citation audit] FAIL
claim: "reduced costs by 43%" — no source, stripping
[Check C — action claims] "I sent the email" → tool result: NOT FOUND
replaced with: "the send did not complete"
[retry] PASS — 3 checks verified, response delivered
That's the whole thesis in six lines. The agent tried to tell the founder something untrue. The rules caught it.
Repo and spec
The spec is on GitHub at wgnr-ai/wOS, Apache-2.0. The README covers the directive table and the enforcement framework. If you're building agents and your answer to "what stops them from fabricating results" is a prompt, this is worth thirty minutes of your time.
Happy 10th anniversary to wgnr.ai, the company this came out of. The next chapter is teaching AI to obey.
Wagner dos Santos is the founder of wgnr.ai, 30 years in brand marketing, now working in AI agent governance. He built wOS and runs the kelle.ai Agent Squad under it in production.
Top comments (1)
The interesting failure is not the first lie, it is the instruction losing priority once the context gets long. A constitution that is checked before delivery is closer to a gate than a prompt. Does the check run on the tool result, or only on the model's summary of that result?
iin1006h11