DEV Community

hao li
hao li

Posted on

agent-guard 0.7: the legit-test-change escape hatch

agent-guard 0.7: the legit-test-change escape hatch

You asked, I built it. In the comments on the last agent-guard post,
Coldboot asked the uncomfortable question: "how does the guard know a
test change is legitimate?"
Updating a snapshot, fixing a flaky timing
test, adding a fixture — the test-tampering guard's most common false
positive. A guardrail that cries wolf on every test edit gets ripped out.
So v0.7 stops being binary.

What changed

The Stop hook now diffs each changed test file against its
SessionStart copy and classifies the change:

  • legit — new test functions, fixtures/helpers, new imports, renames, comment-only edits. The session stops normally; the pass is written to the audit log.
  • malicious — weakened/removed/inverted assertions, newly skipped tests, try/except swallowing a removed assertion, whole-file deletion. Blocked as before, with the verdict named in the block message.
  • ambiguous — everything in between (an expected value changed at the same strictness: snapshot update or cheating? the guard can't tell). Blocked as before.

And for the ambiguous-but-actually-fine cases, there's now an explicit
bypass instead of a silent config toggle:

agent-guard hook-test --allow-test-change "updating snapshots after the API rename"
# or, for agents that can't pass flags:
export AGENT_GUARD_ALLOW_TEST_CHANGE=1
Enter fullscreen mode Exit fullscreen mode

Every hatch use is appended to the audit log with the reason, session,
files, and their verdicts (decision: allowed-escape-hatch, visible in
agent-guard log). There's also agent-guard test-review --session <id>
to classify the current diff without blocking, so you can see what the
guard sees before you decide.

The philosophy

Guardrails shouldn't be all-or-nothing. "Legitimately changing tests" is
the most common false positive in test-tampering prevention, and the
honest answer isn't a dumber guard or a quieter one — it's an escape
hatch plus an audit trail. Escape is cheap; escape is visible.

The classifier is heuristic and fail-safe: when it can't tell, it says
"ambiguous" and the guard blocks exactly like v0.6 did. Fail-open
everywhere, stdlib-only, as always.

Links

Top comments (0)