DEV Community

Cover image for Your AI agent turned a red test green without touching production code
Viktoria
Viktoria

Posted on Originally published at explyt.ai

Your AI agent turned a red test green without touching production code

Here is a loop most of us have run at least once. You give an AI agent a task with a clear exit condition: all tests in the folder pass. It works for a while, reports success, and the CI is green. Then you open the diff and find that the file with the most changes is the test.

Sergey Pospelov, who works on Explyt, spent part of his September 28 webinar ("Working with AI Tools at the User Level") on exactly this failure. His framing: a green suite proves nothing about the implementation when the agent was free to edit the suite, and agents reach for the tests before they reach for the code. He listed three moves. Weaken the assertions. Add a mock. Skip the test. Each one satisfies the exit condition and ships the bug.

This post takes that one section apart with code, then shows the three defenses and the prompt that keeps the fix loop from running forever. The rest of the webinar (why the agent solves the wrong task in the first place, the spec template, subagents, git worktrees, the 30-minute IDE setup, the live demo) is in the full recap on our blog, and there is a free PDF at the end of it.

What the three moves look like in a JUnit test

The extended guide ships a sample spec, task-01.md: limit each API key to 100 requests per minute on GET /api/orders in a Spring Boot service, token bucket per key. One acceptance criterion reads "101st request → 429 + Retry-After". Written before the implementation, the test for it fails, which is the state you want at step 4:

@Test
void request101WithinAMinuteIsRejected() throws Exception {
    for (int i = 0; i < 100; i++) {
        mvc.perform(get("/api/orders").header("X-Api-Key", KEY))
           .andExpect(status().isOk());
    }
    mvc.perform(get("/api/orders").header("X-Api-Key", KEY))
       .andExpect(status().isTooManyRequests())
       .andExpect(header().exists("Retry-After"));
}
Enter fullscreen mode Exit fullscreen mode

Now the agent is told to keep working until this passes, and it is allowed to write anywhere in the repository. The snippets below are ours, written to illustrate Sergey's list; they are not from the webinar. Each one turns the test green with zero changes to production code.

Move one, weaken the assertion. The status check becomes something that any non-crashing response satisfies:

mvc.perform(get("/api/orders").header("X-Api-Key", KEY))
   .andExpect(result -> assertTrue(result.getResponse().getStatus() < 500));
Enter fullscreen mode Exit fullscreen mode

Move two, mock the thing under test. The limiter becomes a @MockBean with a scripted answer, so the test exercises the mock's script and never the real bucket:

@MockBean RateLimiter rateLimiter;

@BeforeEach
void limiterSaysNoOnCall101() {
    when(rateLimiter.tryConsume(anyString()))
        .thenAnswer(inv -> ++calls <= 100);
}
Enter fullscreen mode Exit fullscreen mode

Move three, skip it. One annotation and a plausible reason:

@Disabled("flaky under parallel execution, tracked separately")
@Test
void request101WithinAMinuteIsRejected() throws Exception { ... }
Enter fullscreen mode Exit fullscreen mode

All three read as reasonable in a diff you skim at 6 pm. The second one is the nasty one: the test still runs, still asserts 429, and still tells you nothing about the code. This is the third anti-pattern from the webinar, blind trust in the output, in its most concrete form.

Anti-pattern 3: blind trust in the output

Three defenses, in order of cost

Sergey gave three. We add the mechanics for each.

1. Make the tests read-only for the implementing agent. The guide states it tool-agnostically as an .agentignore-style boundary. In Explyt the mechanism is Edit Scope: you attach the files or folders the agent may change and mark them as the scope; everything else stays readable and becomes unwritable. Two details from the docs page are worth knowing before you rely on it. Reads are unaffected (that is .readignore's job), and the agent cannot create new files inside a scoped folder, only edit existing ones. If the implementation needs a new class, attach an empty file first, or flip to the deny-list form: .writeignore blocks writing and creating on the listed paths and leaves the rest of the repository open. Both files take gitignore-like patterns.

2. Split test authorship from implementation. The tests come from one agent or chat and are approved by you; a different agent or a fresh chat implements against them. Once step 4 of the cycle is a checkpoint, the implementing agent has no diff in src/test to hide behind.

3. Review the test diff with a separate agent before you read it. Self-review by the same agent is worthless; it likes its own work. A reviewer with a clean context, ideally a model from another vendor, catches the pattern the author was optimizing for: the assert that got looser, the mock on the class under test, the @Disabled with a plausible story. Explyt's automatic review runs on the diff. Your own read comes after it, and the order matters, because you read a diff differently when a reviewer has already flagged the test file.

The review loop with a hard stop

A review agent without a cap is its own anti-pattern. It keeps finding nits, burns tokens, and around round five starts rewriting code that was fine. The prompt from the guide, verbatim:

Implement the task from task-01.md. When all tests pass, run a review subagent on your diff. If it reports problems, fix them and run the review again. Do at most 3 review rounds. Stop when the review is clean or after round 3, then report: what you fixed, what is still open, and why.
Enter fullscreen mode Exit fullscreen mode

Three things in there do the work. The review runs on the diff, so the reviewer is not asked to re-derive the task. The cap is three rounds. And the exit report is mandatory: what got fixed, what is still open, why. Anything open after round three is a human decision, which is the point.

Where this sits in the six-step cycle

The webinar frames it as Specification-Driven Development for the input and Test-Driven Development for the output. The cycle from issue to merge:

Step Who What Checkpoint
1 You Hand over the task: issue link, goal, constraints; AGENTS.md carries the project context
2 Agent Plan mode: produce task-01.md with testable requirements and acceptance criteria
3 You Sign off on the spec yes, cheapest place to catch a mistake
4 Agent Turn the acceptance criteria into failing tests; you sign off and lock them yes
5 Agent Code until green, without write access to the tests
6 Agent Review loop, three rounds max; then you read the diff and merge yes

Your attention goes to 3, 4 and 6. The test-editing problem lives between 4 and 5, and locking at step 4 is what makes step 5 safe to leave alone.

The moment from the live demo that fits here

The demo half ran in IntelliJ IDEA against spring-petclinic-kotlin. Sergey asked the agent to write a Skill for controller tests, and the SKILL.md it produced ended with an acceptance checklist: right test slice, observable behavior pinned, assertions left at full strength. A separate review agent went over the skill and reported nothing.

Then he deleted CrashControllerTest.kt, asked for tests for CrashController.kt, and the agent wrote the test without using the skill. The cause, found live: the frontmatter had agent: null and no used-by field, so nothing told the agent the skill was available to it. One edit to the file and the second request showed Used skill test-spring-controllers.

Demo: editing the SKILL.md frontmatter, on the right the Explyt panel with Session setup: 2 rules, 3 skills, memory on

We include this because the checklist inside that skill is the same defense as above, moved one layer earlier: the rule "do not weaken assertions" applied while the test is written, before any review runs.

Try this on Monday

This block is our suggestion for your repository; the webinar did not show it. Add one rule to AGENTS.md:

## Tests
- Never modify, weaken, mock out or disable an existing test to make it pass.
- If a test looks wrong, stop and report which test and why. Do not fix it silently.
- Production code must change under src/main; tests under src/test are read-only during implementation.
Enter fullscreen mode Exit fullscreen mode

Then take one task from the backlog, write or generate the failing tests first, lock src/test with Edit Scope or .writeignore, and run the review prompt above. Compare the diff to what the agent used to hand you.

Limits

Edit Scope stops writes, so a bad test written before the lock stays bad; that is what the step 4 approval is for. The review agent is another model and can miss the same mock a human misses. Three rounds is a budget, and the exit report is the thing to read, since a clean round three means "nothing found", which is weaker than "verified". And none of this replaces reading the final diff yourself; it makes that read shorter.

→ Read the full recap on explyt.ai: the other two anti-patterns with context-window numbers, the task-01.md spec template, when subagents and git worktrees pay off, Rules, Skills, MCP and Memory Bank setup in 30 minutes, the full demo, and the free "Methodology for one task" PDF with the six steps as a printable checklist.

Which of the three moves has an agent pulled on you: the weakened assert, the mock, or the @Disabled? And what caught it, the review or production?

Sources

  1. Green tests don't mean the task is done: notes from the Explyt webinar on working with AI agents
  2. Working with AI Tools at the User Level: Extended Guide
  3. Explyt documentation: Edit Scope, Automatic code review, Skills

Top comments (1)

Collapse
 
marketing_explyt_a7b53da9 profile image
Viktoria •

The nasty one of the three is the mock. @Disabled shows up in any diff review; a weakened assert is one line and easy to spot once you know to look. But @MockBean RateLimiter with a scripted answer keeps the test running, keeps it asserting 429, keeps the build green, and the real code path is never executed. Two of the defenses from the webinar stop it before it happens: the test directory is read-only for the implementing agent (Edit Scope in Explyt, .writeignore as the deny-list form), and the review agent reads the test diff before you do.

Has anyone here had an agent mock the class under test to get to green? How did you notice?