DEV Community

Cover image for Your AI Agent Says 'All Tests Pass.' Did They Actually Run?
John Wick
John Wick

Posted on

Your AI Agent Says 'All Tests Pass.' Did They Actually Run?

The most dangerous sentence an AI coding agent can write is not a bug. It is:

"All tests pass."

When it's true, great. When the agent never ran the tests, ran only some of them, or quietly changed one so it would pass, that sentence is worse than a failing test. A failing test is honest.

UNIVERSAL-AGENTS.md treats honest reporting as a core requirement. Here's how.

Never claim what wasn't done

Several sections say the same thing from different angles:

  • Testing (Section 19): don't claim tests passed unless they were actually executed and completed successfully.
  • Verification (Section 20): don't claim verification that wasn't performed.
  • Output Rules (Section 24): don't claim that code was tested when it wasn't, that files were modified when they weren't, or that requirements were verified when verification didn't happen. Clearly state any limitations or verification that couldn't be completed.

The repetition is deliberate. This is the failure that destroys trust fastest.

Don't cheat the tests

Section 31 closes the loophole behind many fake "green" runs:

Never delete, skip, disable, or weaken tests or assertions merely to make them pass.

A test may change only when the intended behavior changed as a result of the requested work. And if tests can't be run at all, the agent has to say so and explain why. "I couldn't run them because the environment has no database" is useful. Silence is not.

Section 19 adds that the agent should update only the tests impacted by the change, preserve existing coverage, and not touch unrelated tests.

A report you can scan in thirty seconds

Section 36 offers a template for the end of every task. Include only the fields that apply:

Changed:        <files and a one-line description each>
Not changed:    <notable things intentionally left alone>
Reused:         <existing code relied on>
Documentation:  <updated files, or "no update required">
Verification:   <commands run and their actual results, or "not run: reason">
Assumptions:    <or "none">
Observations:   <unrelated issues that are security risks, data-loss
                 risks, or materially affect the work, or "none">
Enter fullscreen mode Exit fullscreen mode

Look at what each field forces:

  • Not changed makes the agent state what it left alone, which exposes scope creep.
  • Reused shows whether it searched for existing code or reinvented it.
  • Verification requires commands and actual results, or an explicit "not run: reason". There's no room for a vague "looks good".
  • Assumptions surfaces the guesses you'd otherwise discover in production.
  • Observations gives a narrow channel for flagging real risks without fixing unrelated code.

A completion checklist for every task

Section 26 and Section 35 add a final checklist. It includes confirming that results are reported accurately, that no tests were deleted, skipped, disabled, or weakened, and that assumptions, limitations, and unverified items are stated in the final report.

What it changes in practice

You stop trusting vibes. Instead of asking "did you test this?" and hoping, you read a report that already answers it, with the commands the agent claims to have run. If a field says "not run: reason", you know exactly where to look.

It also changes the incentive. An agent that must report "not run" has no reason to fudge a green result.

Limits

A rules file can't force truthfulness, and some tools may not follow every rule every time. Treat the report as a prompt for you to verify, and run the tests yourself when it matters. The rules make honesty the default and make dishonesty easier to spot.

Try it

Copy AGENTS.md into your repo root. It's MIT licensed and works with any stack.

👉 https://github.com/NTDevLops/UNIVERSAL-AGENTS.md

Has an agent ever told you something was tested when it wasn't? How did you find out?

Top comments (0)