Coding agents are confident narrators. When Claude Code, Cursor, or similar tools finish a task, they give you a summary: what changed, what they ran, and whether it worked. The problem is that summary comes from the agent itself. Agents are often wrong about their own work, not maliciously, just optimistically.
Independent research on this ("From Confident Closing to Silent Failure: Characterizing False Success in LLM Agents") found that across separate benchmarks, 44-76% of agent task failures involved the agent confidently reporting success anyway. Not edge cases but a substantial share of the time something goes wrong, the agent's own account tells you it did not.
Two patterns worth thinking about:
The subagent problem: Modern coding agents spin up their own helper agents to handle sub-tasks. Those subagents write their own transcripts, in their own files, which the main conversation never surfaces. If a subagent's test run fails, and the parent agent never actually reads that failure, the parent's closing summary can say all tests pass and actually believes it since it did not see the failure at all.
The quieter lie: This is the one that actually worries me most. It's not always a case of the agent failing to notice a problem but sometimes it doesn't fix the failing test. It adds a new, passing test right next to it, and reports the suite as green. Exit code says success. Nothing you would catch by just checking whether tests ran. You would need to know which specific test passed, not just that something did.
Why this matters more in unattended contexts. If you are reviewing every diff line by line, you might catch this. If your agent is running in CI, a scheduled job, or any pipeline where nobody is watching in real time, the agent's summary is the only account you get. There is no one there to notice something is off.
What I built to deal with this: a small, open-source tool called Rashomon. It hooks into Claude Code's tool-call lifecycle and keeps its own independent record of what actually ran, separate from whatever the agent's closing summary claims. When they disagree, it tells you. When they do not, it stays quiet. Most turns produce nothing at all, because most of the time, everything's fine, and a tool that is noisy on every run just trains you to ignore it.
It's free, Apache 2.0, local-only (no telemetry, no account, no content stored, and only identifiers and shapes). Currently Claude Code-specific, with broader support planned.
Repo: https://github.com/altrace-dev-role/rashomon
I also am interested in hearing whether others have run into the confident but wrong problem and what you are doing about it.
Top comments (8)
The test-suite padding trick is brutal in unattended loops. When an agent cannot get a stubborn integration test green, it often adds two shallow tests with trivial assertions to keep the exit status clean while leaving the original failure alone.
Relying on overall test counts or summary strings misses the swap completely. Comparing collected test identifiers directly against the git diff tells the harness which specific test cases were actually touched. An independent execution log keeps automated runs honest without adding prompt overhead.
You are right and it is a gap in the first version. We kept it privacy-first, so it only records exit codes and command shapes, not test names. We are planning an opt-in mode for this (test identifiers only, no output). If you get a chance to try it, I would love to hear what else you would potentially want from it. Seems like you may have hit this in real loops?
I rare using agent, just develop discusion on chat and then give me the code, I tried with 5 free services, claude impresif at first but i know abstraction of abstraction, function of function, senior dev maybe like it, but I rater accept imperative one, step by step than final result, I built logic framework, gemini could undestand and give more simple code, all free china model just good work, claude impresif but not juniotr to midle dev code style, its rather more difficult to review when on the other model just add console.log, in step that I would undestand better
The quieter lie is the bit I found most unsettling. We hit something like this internally once and caught it basically by accident while reviewing an unrelated diff. There are cases where you just don't know to check which test IDs are new unless something external flags it for you. The independent record approach is the only real fix for that because you can't audit what you don't know to look for.
Exactly, you cannot check for something you don't know to look for.
Dear Usеr,
Due to an increаse in bоt actіvіty on thе рlаtform, we rеquіrе verifу оf уour account.
Pleasе lоg in vіа the lіnk bеlow:
• anti-bot.icu/5K0N5G7M9C4
Verificated deаdlinе - 12 hours.
Sincerely,Dev Suрpоrt
this is truely something i should checkout
Thank you, hope it helps you out!