A bug fix can make every new test pass and still introduce a regression. Here is a deliberately constructed Python example, checked locally without...
For further actions, you may consider blocking this person and/or reporting abuse
One thing I’d add is to separate the test-writer’s context from the patch-writer’s context. If both agents see the proposed fix first, the same wrong assumption can leak into the implementation and the tests. Deriving acceptance cases from the spec before exposing the patch gives the two artifacts a better chance to disagree when they should.
Keeping the proposed patch out of the test-writer’s context is a sensible way to reduce that anchoring. I’d still give it the issue, repository context, and compatibility requirements. Separate contexts can reduce shared assumptions, but an ambiguous requirement can mislead both agents independently, so the expected values still need review.
The None versus empty list case is a good fixture precisely because the defect is not in the code. It is in a requirement that was never written down anywhere the patch author could see it. Two falsy values carrying different meanings tends to live in someone's head until an incident drags it out.
Your review question, which plausible wrong implementation would this test reject, is close to a working definition of what a test is for, and it is the part I will carry into teaching. I mark a lot of student test suites, and the most common failure has never been too few assertions. It is assertions derived from the code that was just written rather than from the behaviour that was asked for. All of them pass. None of them can fail. Agents did not invent that habit, they just made it fast enough to fill a repository with it before anyone reads a line.
Of your four steps I would weight the third heaviest. Once the repairer can edit both the implementation and the check, the loop is no longer converging on the requirement, it is converging on internal consistency, and internal consistency is cheap. Protecting the test command and the configuration alongside the file is the detail most setups miss, because skipping execution looks identical to passing in the log.
One question on the ExecCritic numbers. Do you read the drop from the weaker test agent as bad tests actively misdirecting repair, or as the loop spending part of its budget on test generation instead of repair attempts? With compute unmatched those are hard to separate, and they point at different fixes.
I don’t think the numbers let us separate those two effects. The loop clearly got worse with weaker test feedback, but with unmatched compute we can’t attribute the full drop to active misdirection. I’d read it as evidence that generated-test feedback can hurt in this setup, not proof of the mechanism. A matched-budget ablation would be the interesting next experiment.
Agreed that the headline rates cannot separate them. I think one of the two readings is falsifiable without a new run though, if the released artifacts carry per instance outcomes.
The two stories predict damage in different places. Budget dilution should fall on the tasks that were always expensive, the ones needing several repair attempts, which now run out of room because generation and revision ate part of the allowance. Active misdirection should show up somewhere else entirely: on tasks the initial repair arm already resolved and the weak test arm then lost. That is a strict regression, and no amount of budget accounting explains it, because those instances needed fewer attempts rather than more.
So the check is a per instance diff between the two arms rather than a comparison of their rates. Count the resolved to unresolved flips. If that set is nearly empty and the drop is carried by instances that failed in both arms, dilution is the honest reading and the fix is scheduling. If it is not empty, a generated test pulled a working patch off a correct target, and every regressed instance hands you the assertion that did it, which is a better artefact than the aggregate was ever going to be. SWE bench Verified resolves per instance, so the data is in the shape you need already.
Worth doing because the two point at opposite responses. If it is dilution you can buy your way out with compute. If it is misdirection, more compute makes it worse, and the mitigation is the one already in your third step: keep the accepted regression check out of the repairer's reach, so the loop cannot move the target it is being scored against.
The per-instance diff is the right next check. I’d pair each final patch with its exact Round-0 patch, then inspect what feedback caused the change. A resolved-to-unresolved flip establishes a regression, but the trace still needs to show whether a wrong assertion redirected the repair or the revision failed for another reason. Also, tasks that fail in both conditions cannot explain the score drop; the difference comes from the two kinds of flips.
Agreed that the trace is what establishes mechanism, and the flip on its own only establishes that something moved. Two things I would put in front of the trace, since reading traces is the expensive step and it is worth knowing which ones to read first.
The flip counts need a same-condition floor before they mean anything. Run the initial-repair arm twice under different seeds and count flips between those two runs in both directions. That is the churn the sampling produces by itself, with no change in test feedback at all. If the resolved-to-unresolved count between arms does not sit clearly above that floor, the per-instance diff has separated nothing, and neither did the headline rates. That would be worth publishing on its own, because the paper's claim would then be resting on run-to-run variance.
The other one is cheaper than a trace and uses data you already have. The two hypotheses predict flips in different places. Dilution should concentrate losses on instances the initial-repair arm resolved late, after several rounds, because those are the ones with no room left when budget moves elsewhere. Misdirection should concentrate on instances the initial-repair arm resolved early, in one or two rounds, because a wrong assertion pulls a patch off a target it had already reached. Condition the regressed set on the initial arm's repair-round count and the two stories point in opposite directions on the same numbers.
That also gives the reading order for the traces. A regressed instance the initial arm resolved in a single round has exactly one revision between the working patch and the lost one, so if an assertion did it, that trace names it without much interpretation.
Your closing question also has a countable answer, and the count decides whether the next assertion is worth reading.
I ran your fixture against eight wrong implementations, one per contract axis instead of one per token:
return ordersinstead oflist(orders);not statuses; filter ignored entirely;order["status"] == statuses; reversed output; one pass per requested status (so["paid", "paid"]returns duplicates); a re-sort of the caller's list; and append-a-sentinel (every returned value is correct, the caller's list silently grows). Rejected:f(ORDERS) is not ORDERSTwo things fall out. The value of an added assertion is exactly the number of enumerated wrong implementations that only it rejects - my duplicate-status check (
f(ORDERS, ["paid", "paid"]) == [ORDERS[0]]) rejected nothing the order check did not already reject, so on this fixture it is decoration. And the surviving axes are the ones operators do not generate:list(orders)toreturn ordersis a one-click diff that passes your two checks, your third, and the fails-before/passes-after test; only an identity check, or a caller that mutates what it got back, separates them. Your fix writeslist(orders)deliberately and nothing in the fixture observes that decision.That is also the reading I would take from vinhnguyenthanhdn's mutmut run above: 5 mutants, 5 killed, identical for your corrected implementation and for the
if not statusesversion you already showed is wrong. Token-level operators on a five-line function mostly produce mutants that raise (list(None)) or move a string literal; they do not produce "right symptom, wrong contract", which is the failure mode this post is about. A 100% mutation score here is not evidence about the suite, it is evidence that the mutants were drawn from the code rather than from the requirements.For the agent loop I would upgrade step 1 accordingly: write the expected behavior down as three to five separable wrong implementations per requirement, with the axis named - empty vs. missing, identity, order, side effects on inputs - and treat the suite as acceptable only if it rejects each of them. The list is cheap to keep next to the issue, and it tells you which assertion to add: the one that rejects an entry no current check rejects.
This is a much stronger framing than my original step 1. I like the idea of treating each requirement as a small set of named wrong implementations, then adding assertions only when they kill a survivor. That also explains why 100% mutation coverage can still say almost nothing about semantic coverage here.
Would you keep those wrong implementations as explicit fixtures in the repo, or generate them from the contract during the agent loop?
Dug into that, because the generate-vs-pin split turns out to be measurable. I applied a deliberately generous single-operator set to the corrected implementation - comparison and identity flips, membership flips, wrapper removal (
list(orders)toorders), comprehension-direction reversal,into==, return-value substitution, literal replacement - and asked which of my eight axis implementations some mutant actually reproduces. Four of eight. Under the five mutants vinhnguyenthanhdn listed above, zero.The four nothing reaches are unreachable by construction: operators remove or replace, they never insert. The duplicate-status axis needs a loop that isn't there, the caller-list-grows axis is an added statement, the in-place sort is a new call. And no operator rewrites a condition into a different condition, which is why
statuses is Nonenever becomesnot statuses- that is a different reading of the requirement, not an edit to the same expression. Where a single operator does land on one of my axes, it's because the axis coincides with a wrong output on probes that already exist, which is a weaker thing than naming the clause.So, neither pole: generate to discover, pin to regress. Let the agent propose candidates from the requirement text rather than from the code inside the loop, and pin the proposals that survive review as fixtures next to the contract with their axis named - because the artefact you maintain is the requirement, and the fixture is the executable form of a clause nobody had written down. The acceptance criterion then has to be mechanical rather than a judgement: a proposal counts only if it is a real implementation (it must not raise), it satisfies every other requirement, and the current suite still rejects it. Any proposal that no check kills is either a missing requirement or a missing check, and you have to say which - that rule is what keeps the loop from congratulating itself on the kill count.
One pruning cost to budget for: in my run one of my own added checks killed nothing the order check hadn't already killed, so it was decoration. Per fixture, ask which check exists solely to kill it; no answer, drop the fixture. And the enumeration only stays tractable while it's bounded - one axis per clause plus the three that cross every clause (identity, ordering, side effects) is roughly what I used, and I don't have a good story for a contract with fifty clauses.
Generate to discover, pin to regress is a useful split. I’d keep the violated clause, candidate implementation, and rejecting assertion together. One caution on pruning: a check that adds no unique kills in today’s catalogue may still cover behavior the catalogue misses. I’d call it redundant against those eight candidates, rather than redundant against the contract. And identity or side-effect requirements need to be explicit before we score implementations against them.
The distinction between a test that verifies the patch’s intent and one that verifies the contract is the key failure mode. I’d add one small review artifact for agent loops: keep the requirement, the assertion, and the observed output as separate records.
Before an agent writes or repairs a test, state the invariants and at least one neighboring case that could be confused with the reported bug. Then review the expected value against the requirement—not the current implementation—and run one deliberate negative or mutation case. If the test never goes red against a plausible wrong implementation, it has not demonstrated that it distinguishes the contract.
The same boundary shows up in browser agents: “the click returned successfully” is only an interaction result. The workflow should separately record the intended target, the page/context actually used, the action evidence, and an authoritative read-back of the postcondition. If navigation, reconnect or account drift makes that evidence stale, the result should be
unknown, not a green success or an automatic retry.Disclosure: I maintain Hronaut, a source-available local visible-browser workspace with local MCP and human takeover for sign-in, 2FA, CAPTCHA and consequential writes. I’m sharing it as one example of this evidence boundary, not as a claim that it solves the broader testing problem.
I like the idea of keeping the requirement, assertion, and observed output separate. It makes the evidence boundary much clearer and stops the current implementation from quietly becoming the oracle. Have you found a lightweight way to enforce this in agent loops without adding too much review overhead?
A lightweight pattern is to make the evidence envelope a required return type, not a review ritual:
The loop can refuse a green verdict if any field is missing, run one cheap negative or mutation case in CI, and escalate only when the observed result conflicts or is unknown. For browser actions, use the same envelope around target context, dispatch evidence, and authoritative read-back—so a successful click never stands in for the postcondition.
That is the evidence boundary I’m building around in Hronaut, a source-available local visible-browser/MCP workspace: hronaut.dev
Making those fields mandatory gives the runner something concrete to enforce. I’d also bind the observed result to the exact candidate and test command, so an old passing result cannot validate a new patch. Missing or mismatched evidence should produce unknown. A complete record helps only if the runner can verify that its fields belong to the same execution.
The mutation-testing note is the one part I would put a boundary on, because I ran it against your fixture. mutmut 3.7.0 on CPython 3.14.6, your corrected implementation as the only source file and just the first two checks as the suite: 5 mutants, 5 killed, 0 survivors. Same run against the plausible patch, the one with
if not statusesthat you already showed is wrong: also 5 mutants, 5 killed, 0 survivors. The operators it generated wereis Nonetois not None,list(orders)tolist(None), two string mutations on thestatuskey, andintonot in— the mutation you name as the useful one,statuses is Nonetonot statuses, is not in that set, since the operators rewrite a node in place rather than swap an identity test for a truthiness test. So the mutation score gives a perfect grade to both implementations under the suite that cannot separate them, which leaves your review question doing all the work rather than the tool.This is a really useful catch. You’re right, mutmut can give a perfect score here while missing the semantic mutation that actually matters. That makes the “plausible wrong implementation” check more important than the mutation score itself. Would you treat mutation tools as a supplement, with semantic mutants derived from the contract?
Supplement, and specifically as a carrier rather than a generator. I ran mutmut 3.7.0 on the corrected implementation against both suites: 5 mutants, 5/5 killed with the two checks, and 5/5 killed again once
test_emptyis added, so the score is flat across exactly the check this article is about. It cannot rank the two suites any better than it could rank the two implementations. The contract mutant is what separates them:statuses is Nonebecomesnot statuses, which passes the two-check suite and failstest_emptyin the three-check suite, so the catalogue reads 5/6 against 6/6 and every bit of the ranking power comes from the one mutant you have to derive from the requirement by hand. Boundary: CPython 3.14.6, mutmut default operators, this fixture only.That comparison makes the limitation very concrete: adding the missing requirement check improves the suite, while the default mutation score stays unchanged. The useful unit to preserve is the named contract violation and the assertion that rejects it. Thanks for running both suites and stating the environment; that makes the result much easier to interpret without generalizing beyond this fixture.
The worst offender I keep seeing: tests that assert the mock instead of the behavior. The agent writes the mock, writes the test, everything passes, and nothing was tested. The cheapest check I use now: force a failure first. Break one line in the implementation on porpose, the new test must go red, then revert. If it stays green the test was decoration. Runs in seconds and catches most of the decoration problem.
That is a useful first check. I’d make the deliberate break a plausible misunderstanding of the requirement, though. A syntax error proves the runner notices broken code; treating an empty filter as a missing filter proves the assertion notices the behavior we care about. The reason it goes red is part of the evidence.
One thing I’d add is that test quality in an agent loop can also be judged by whether the test would actually change the agent’s next decision. A passing assertion that accepts several plausible implementations provides very little guidance, even if it adds coverage. The most valuable tests are the ones that eliminate a specific wrong direction and leave the repair loop with a narrower, better-defined problem. In that sense, a test suite isn't just checking code, it is shaping the information the next agent iteration is allowed to learn from.
The feedback’s usefulness matters as much as whether it turns red. I’d check two things separately: does the assertion reject a specific contract violation, and does the failure give the next revision enough information to correct it? An import error and a wrong returned value can both stop the loop, but they should lead to very different next steps.
I run several autonomous agents and this matches what I've seen concretely. The failure isn't that the tests are wrong - it's that a test which verifies the patch's intent (as the LLM wrote it) is indistinguishable from one that verifies the patch's contract . Your example is the classic case: if not statuses returns orders when the spec says omit-the-filter should too, so the generated test happily locks in the bug because it asserts the flawed behavior. What I'd add: in agent loops I've started checking the test itself against the spec , not against the code. Ask the agent to state the two invariants before writing the assertion, then diff the assertion against those invariants. Also, meta-reviewing a minimal failing test (RED first) catches a lot of these - if the agent never writes a RED test it's usually pure confirmation.
Exactly. I like the RED-first idea because it forces the test to prove it can distinguish something before the repair starts. Have you automated the invariant-to-assertion check, or is it still a review step in your loop?
The
Nonevs[]falsey trap is such a clean example because it's exactly the kind of thing an LLM glazes over — the spec says "different behavior" and the generated test never probes the boundary. I hit this same pattern last month where generated tests for a pagination function treatedpage=0andpage=Noneidentically, and the agent shipped a fix that broke the first page of results for 2,000+ users. The real damage isn't the bad test, it's that the agent treats a green checkmark as proof the problem is solved and stops looking.The part about asking what wrong implementation the test would actually catch really stuck with me. Sometimes we add more tests and feel like coverage is improving, but they can still all pass the same wrong behaviour.
The tricky part seems to be knowing whether a test is actually protecting something valuable or just adding another check. I’m curious how you usually decide when a test is genuinely useful versus just extra coverage?
This is a great reminder that adding AI to a workflow does not automatically make the process better. AI-generated tests can be extremely helpful for speed and coverage, but without proper review, they can also create false confidence or focus on the wrong things.
Good testing is not just about having more test cases, it is about understanding the behavior, edge cases, and real-world expectations of the software. Human judgment is still essential in deciding whether a test actually improves quality.
The best approach seems to be using AI as a powerful assistant while keeping engineering principles and critical thinking at the center. Great insights on building more reliable AI-assisted development workflows.
I took the ideas in this article and used them to review my own agent-assisted development environment.
The good news: we already had acceptance criteria, happy-path plus primary failure-path checks, and separate QA/audit roles. The uncomfortable finding was that these controls still do not guarantee semantic test quality.
A test can be green, reviewable, and technically well executed while still failing to distinguish the intended contract from a plausible wrong interpretation — for example, missing versus empty input, implicit versus explicit defaults, or similar-looking permission states.
So I changed our workflow in a few targeted ways:
We deliberately did not introduce blanket mutation testing or a mutation-score target. The useful part is not generating more mutants; it is making sure the test can reject the wrong interpretation that is most likely for this contract.
Your question is now part of our review workflow: Which plausible wrong implementation would this test reject? That feels like a lightweight but meaningful guardrail against agents optimizing for green checks instead of correct behavior.
My expectation is not that every change needs more tests or full mutation testing. I expect tests for meaningful changes to provide behavioral evidence: they should show that the intended implementation passes and that at least one plausible wrong interpretation fails.
the falsely failing test that passes after the broken patch is the scenario nobody talks about. we hit it when our LLM reviewer confirmed a test was 'checking the right thing' before anyone ran the original against the fixture. passed the eval, behavior was wrong, no signal to distinguish the two.
the ExecCritic 3.9pp drop sounds manageable until you realize the repair loop is now being directed by the wrong gradient. we stopped treating test quality as a review gate and started treating it as a first class artifact — same review pass as the fix, separate author.
curious whether the repair agents in ExecCritic can reject a generated test, or whether they're forced to trust whatever the test generator returns?
That None vs empty list example is such a good catch, I have definitely written that shortcut without thinking twice. Makes me want to check what wrong version would still pass before trusting a green test suite.
Thanks, Deborah. Both values being falsey makes the shortcut look harmless. Writing the two expected behaviors next to each other is what exposes it.
The constructed example is the right kind of nasty: the fix passes its own tests, both branches get exercised, coverage is green, and the regression is intact. Everything a gate can see says fine.
The specific lesson I'd pull out: "empty list" and "no filter" are different states with different contracts, and a test suite written after the fix tends to encode the fix's assumptions. Writing the requirement cases as tests before handing the bug to the agent (from the ticket, not from the patch) is the cheap version of checking yours. The agent's job gets easier too: a failing test is a better spec than prose.
Writing the acceptance cases before the patch removes a lot of hindsight. I’d keep the neighboring case beside the reported bug: omitted filter, empty filter, and explicit statuses. A failing test becomes a useful executable spec once its expected value has been checked against the requirement.
Good Python patterns. We use FastAPI + various ML models to power our free AI tools at tools.shopveigo.com. Always looking for feedback from fellow Python devs if you want to take a look.