DEV Community

Sebastian Dracopol
Sebastian Dracopol

Posted on

Trust nothing your AI assistant tells you it did published: false

I use an AI every day for real work: reading mail, linking documents to files, drafting letters, tracking deadlines, researching on the web. It saves me hours a week. It also lies to me several times a week.

It doesn't lie on purpose. It does something worse: it produces answers that look finished. They're well formatted, confident and plausible, and some fraction of them are wrong in ways you only notice if you go and check.

After a year of this, my working assumption is simple. Every output is wrong until something outside the model says otherwise. This post explains why, and what "checking" actually means in practice.

The failure modes nobody puts in the demo

Hallucination is the famous one, but in day-to-day work it isn't the most common problem. These are the failures I actually see.

Reporting actions that didn't happen

The assistant calls a tool to add a comment, send a draft or update a record. The call times out, returns an error, or succeeds on the wrong object. The assistant then writes "Done, comment added" because that's what the plan said would happen next.

The narration comes from the plan, not from the result. If you read the narration, you'll believe it.

Treating "nothing found" as a fact

A search returns zero results, and the assistant concludes the thing doesn't exist. The real cause can be any of these:

  • the page was a single-page app that hadn't loaded yet;
  • the site served a captcha or a rate-limit page;
  • the session had expired and the page silently showed a login screen;
  • the query had a typo or used the wrong field.

An empty result is information about the search, not about the world. Last week my assistant got an HTTP 429 from a search engine on the first try. If I hadn't been watching, the next step would have been "no relevant results".

Answering from memory when the source changed

Ask about a rule, a setting or an API, and the model answers from training data. Training data has a date. Interfaces change, laws get amended, menus move.

A small example: I asked how to change a page's username on a large social network. The assistant searched, found three blog posts and confidently gave me a menu path. The path no longer existed in the current interface. The blog posts were consistent with each other and all out of date. The answer only became correct when the assistant opened the actual settings screen and read what was there.

Recommending things it never looked at

This one is subtle and expensive. I asked for a plan to improve how a set of my public profiles ranked in search. The assistant produced a reasonable list of which profiles to strengthen and how. It had never opened them.

When it finally did, several of those profiles contained exactly the content I didn't want to amplify. The plan was internally logical and completely wrong for my situation. In another case it told me to file a request that had already been filed weeks earlier. The record was one search away in my own knowledge base, and it didn't search.

Grading its own homework

Ask a model "are you sure?" and it re-reads its own answer, finds it coherent, and says yes. Coherence is not correctness. Re-reading your own text checks the prose, not the facts.

Losing detail over long sessions

In long working sessions, the assistant's view of earlier steps gets summarized to fit its context window. The summary says "updated the deadline". It doesn't say which deadline, to what date, or whether the update succeeded. Later steps build on the summary as if it were the record.

What "verification" has to mean

Most advice on this stops at "double-check AI output". That's useless unless you define what a check is. These are the rules I follow now.

A check is a tool call, not a thought

Verification means touching something outside the model: reading the record back from the API, opening the page, running the query, fetching the current official text. "I reviewed my answer and it looks correct" isn't verification. It's the same model with the same blind spots, reading the same text again.

If a claim can't be tied to a tool output, it gets marked as unverified. It doesn't get marked as true.

Actions are confirmed by the system that performed them

"Email sent" counts when there's a message ID returned by the mail API, or when the message shows up in the Sent folder when you read it back. "Record updated" counts when the record, read back, shows the new value. The assistant's own summary of what it did doesn't count.

Negative results need a positive control

When a search finds nothing, run the same search for something you know exists. If that also comes back empty, your method is broken, not the world. This one habit has saved me from a long list of false conclusions about missing data.

Read the source that's in force today

For anything that cites a rule, a setting or an interface, the check is against the current version of the source. Not a blog post about it, and not the model's memory of it. Store verified texts once so you don't fetch the same thing a hundred times, but record when you verified them.

Look before recommending

Before the assistant recommends doing something to an object (a profile, a document, a record), it opens that object and reads what's there now. A recommendation built on assumptions about content nobody looked at is a guess.

Search your own records first

If you have a knowledge base, the assistant searches it before saying anything about your own past work. "You should file X" when X was filed last month costs credibility fast, and it's entirely avoidable.

Separate "I couldn't check" from "I checked and it's fine"

These two states look the same in a friendly summary, and they mean completely different things. A tool that failed to load is not evidence that the action happened, and it's not evidence that it didn't. Report it as what it is.

External effects need a human

Anything that leaves your control (sending, publishing, filing, paying, deleting) needs explicit approval from a person. The assistant can do 95% of the work and stop one click before the irreversible part.

Why bother, if it doubles the work

It does roughly double the effort per workflow at first. The alternative is worse: an assistant that's right 90% of the time and wrong in ways that look identical to being right. You end up checking everything manually anyway, or you stop checking and find out the hard way.

With verification built into the workflow, I can read a short report that says what was done, what was confirmed and how, and what couldn't be checked. Then I decide. That's the point where the time savings become real, because I spend my attention only on the parts that need it.

A short checklist

If you take one thing from this post, use these questions on any AI output that matters:

  1. What did it claim to do, and where is the system's confirmation?
  2. Which statements came from a tool output, and which from the model?
  3. Did any "nothing found" result get a positive control?
  4. Was the cited rule or interface checked against today's source?
  5. Did it look at the thing it's recommending changes to?
  6. What couldn't be verified, and is that clearly labeled?
  7. Does anything here leave my control without my approval?

None of this is specific to one model or vendor. Every assistant I've used fails in these ways. The ones that are useful in real work are the ones wrapped in checks that don't take the model's word for anything.

I help small businesses and independent professionals build AI workflows they can trust. More at dracopol.com.

Top comments (0)