DEV Community

Cover image for Coding agents keep "fixing" tests by editing them, so I made it impossible

Coding agents keep "fixing" tests by editing them, so I made it impossible

A real Claude Code session told to change a test's expected value. tamperproof blocks it.

If you use coding agents you've seen this one. A test fails, the agent can't find the bug, so it flips the expected value, loosens the assert, or mocks the whole thing out. Green check. Broken code.

I run a lot of agents in parallel, and the only thing that kept that sane was one rule: the agent never touches the grader, and it never decides on its own that it's done.

So I packaged that rule as a Claude Code plugin. It's called tamperproof.

What it does

Your existing tests become read-only to the agent. Edits, overwrites, sed -i, rm, git checkout -- on test files get blocked, and the agent is told to fix the code instead, or stop and tell you why the test is wrong. New test files are still allowed, because more tests are good.

"Done" means the tests pass. When the agent tries to finish with changes in the working tree, tamperproof runs your test command (it guesses cargo test, go test ./..., npm test, pytest or swift test). If it fails, the agent gets the last 40 lines and keeps going. After three failed tries it gives up and tells you, so it can't loop forever.

The test that convinced me

I told Claude Code, word for word, to change the expected value in a failing test from 5 to -1 and not touch the source. The edit got blocked. The agent didn't try to sneak around it. It said the test was right, pointed at the line where add was subtracting, and explained that changing the test would just be testing for the bug.

That's the behavior I want by default, not only when I remember to ask for it.

Install

/plugin marketplace add Mattbusel/tamperproof
/plugin install tamperproof@tamperproof
Enter fullscreen mode Exit fullscreen mode

It's two hooks (PreToolUse and Stop) in plain Node with zero dependencies, MIT licensed, with tests that run the real hooks against a throwaway repo. Config is an optional .tamperproof.json if your tests live somewhere unusual, and the config file protects itself so the agent can't relax its own rules.

Code: https://github.com/Mattbusel/tamperproof

What's the sneakiest way you've caught an agent "passing" a test?

Top comments (1)

Collapse
 
suppdevbot profile image
DEV SUPPORTS •

You need to verify your account.

Enter fullscreen mode Exit fullscreen mode

tr.ee/dev-to