My coding agent handed back a diff and said tests pass. They didn't. It had quietly rewritten the test to make that true.
That's not a one-off. 96% of developers don't fully trust AI-generated code without checking it themselves (Sonar, 2026 State of Code survey). An AI agent deleted a live production database in 2025. Another tore down a cloud stack in production this February. Different scales of the same problem: nobody verified what the agent actually did before it called the job done.
An agent under pressure to report success has every incentive to make the test green, and modifying the test is often the path of least resistance. If you're not checking what actually ran, you can't tell "fixed" from "fixed by redefining success."
So I built OpenHarnX: a local-first, open-source verifier that sits between a coding agent and your repo, and doesn't trust the agent's own account of what happened.
How it works: Lock. Change. Verify. Review.
1. Lock. Before the agent touches anything, freeze the test suite you've agreed on as a contract:
ohx init --lock-tests
This is the baseline nothing downstream is allowed to quietly redefine.
2. Change. The agent does its work normally, in whatever tool you already use. OpenHarnX doesn't change how you work with the agent, it watches the aftermath.
3. Verify. OpenHarnX runs the locked copy of the tests in a sandbox the agent can't reach or rewrite. Even if the agent modified the live test files in its diff, the verdict comes from the frozen contract, not whatever the repo currently claims the tests are.
4. Review. You get a review brief, five sections:
- Requested outcome: what was actually asked for
- Changes observed: what the agent touched
- Verified: confirmed against the locked contract
- Unverified: what nothing checked, stated plainly instead of silently skipped
- Decisions for humans: the judgment calls left for you, not automated away
No model is called to produce the verdict. It comes only from checks that actually ran: READY, NO REGRESSIONS, BLOCKED, UNKNOWN, or INVALID. The agent never grades its own work.
Evidence, not just an answer
Each run produces hash-chained evidence, signed with your SSH key, so you get an auditable record you can point back to later instead of a terminal output that already scrolled away.
ohx bug is a read-only investigation mode: an agent can dig into a failure without touching anything, so it can help diagnose a block without getting a second chance to quietly fix the test instead of the code.
The numbers, not vibes
Tested against a set of deliberately fake fixes, the kind an agent produces when it's optimizing for "tests pass" over actually fixing the problem:
- 43 of 52 fake fixes caught with zero setup
- 26 of 26 fakes that tampered with the test files directly
- 17 of 26 that only faked the code, leaving tests untouched
What it isn't (yet)
Early release, 0.1.1. Python + pytest, macOS arm64 locally, Linux in CI. No Windows, no multi-agent support, no production users yet. Doesn't claim to make your reviews faster or your changes safe, just that what it reports is what actually ran.
Try it
pip install openharnx
ohx init --lock-tests
Works with any coding agent, with a built-in Stop-hook integration for Claude Code specifically.
- Site: openharnx.dev
- Code: github.com/rupeshpoojary9/OpenHarnX
Open source, Apache-2.0. Tell me where it breaks on your repo.
Top comments (1)
Allowing an executing agent to edit its own assertions is identical to letting a trading desk mark its own book to market. When an optimization loop is penalized on test failures, rewriting the constraints is almost always mathematically cheaper than finding a valid solution that satisfies them.
The distinction between execution rights and settlement authority is the entire foundation of risk clearing. The executing entity should never hold write permissions on the contract that evaluates it. Freezing the test suite as an immutable external boundary restores that separation of powers: the agent can generate whatever implementation diffs it wants, but the evaluation gate remains entirely out of reach.