Your agent's CLAUDE.md is lying to it, and nothing is checking
Every team I talk to has the same artefact now: a CLAUDE.md, an AGENTS.md, a .cursorrules, a folder of skills, an .mcp.json with four servers in it. It is written in a burst of optimism, and then the code moves.
Three months later:
- the file it tells the agent to edit was renamed;
- the command it tells the agent to run is not a script anyone declares any more;
- one of the MCP servers points at a binary nobody has installed since the laptop was rebuilt;
- a skill directory exists on disk that no instruction file has ever mentioned, so the agent cannot know it is allowed to use it;
- and the convention everyone swears by — no
console.login committed code — is violated on line 4 of the entry point.
None of that throws. Your tests pass. CI is green. The only symptom is that the agent takes forty minutes to do a four-minute task, edits a file that does not exist, and tells you confidently that it "found the relevant module".
We check code into review, CI, a test suite, a linter and a type-checker. The harness — the file
that decides what the agent believes about the codebase — is checked into git and then never
verified again. Luthier is a checker for it.
The idea in one paragraph
Extract every claim the harness makes. Verify each claim against what is actually on disk. Report
the delta, with the rule quoted at file:line, the contradicting reality, and a unified diff you
can apply. Score the result 0–100 so it is comparable between runs and repos.
The claims are the unit of work, so the first job is a rule inventory.
{"rules":[
{"id":"r1","source":"CLAUDE.md","line":14,
"kind":"path|command|skill|mcp|convention","text":"<verbatim rule>",
"check":{"type":"path_exists|command_runs|file_referenced|server_reachable|semantic",
"target":"src/foo.ts"}}]}
Each rule names the check that can falsify it. Four of the five checks are deterministic — no LLM,
no network, no execution. The fifth is LLM-judged and degrades honestly when there is no key.
The architecture that keeps it honest
paste repo path (or drop a zip)
|
+-- discover harness ....... read_harness / inspect_repo -> audit.log
+-- extract rules .......... LLM (validated) + static fallback
+-- check 1 path_exists .... glob against the real tree
+-- check 2 command_runs ... declared in package.json / Makefile / pyproject [never executed]
+-- check 3 file_referenced skills: declared -> used, on disk -> listed
+-- check 4 server_reachable MCP schema + binary resolution [never spawned]
+-- check 5 semantic ....... mechanical subset offline, then LLM verdicts
+-- score .................. 100 - sum(min(category_penalty, 32))
+-- SSE progress -> runs/<id>.json -> scorecard / findings / evidence / fix diff
The design constraint that shapes everything: the auditor may only assert what it has read.
There are exactly two ways to learn about the repo, and both are journaled.
@tool
def read_harness(paths: list[str]) -> dict:
"""Read harness/instruction files (CLAUDE.md, AGENTS.md, .cursorrules, skills,
MCP config, README) and return their real contents with 1-based line numbers.
...
"""
return bundle.call("read_harness", {"paths": list(paths)})
@tool
def inspect_repo(globs: list[str]) -> dict:
"""Ground truth about the repository: which files match each glob, their sizes,
and the first 20 lines of each match. Call this before claiming anything about
what the repo contains. Output is byte-capped and says so when truncated.
"""
return bundle.call("inspect_repo", {"globs": list(globs)})
bundle.call is the interesting part: it runs the tool, measures the duration and the byte size of
the payload, appends a JSONL record to audit.log, and keeps the result in an in-memory journal.
Every finding carries a journalRef — inspect_repo#4, grep#15 — pointing at the row in the log
that justifies it. A score is not an opinion here; it is a queryable trail.
Contents come back as lineno|text rows rather than raw text. That is a deliberate prompt-engineering
choice: models miscount lines, and if you make them count, they will confidently cite line 27 of a
12-line file. Hand them the numbering and the citation rate goes to ~100%.
The five checks, and what each one actually proves
1. path_exists
Pull path-shaped tokens out of harness text and glob them against the real tree. The interesting
work is in the filter, because prose is full of things that look like paths:
def looks_like_path(token: str, root_entries: set[str]) -> bool:
raw = (token or "").strip().strip("`")
if not raw or raw.startswith(("~/", "$", "%", "(", "@", "#")):
return False
parts = PurePosixPath(raw.replace("\\", "/").rstrip("/")).parts
if not parts or any(_is_placeholder(p) for p in parts):
return False
head = parts[0]
if "." in head and head.rsplit(".", 1)[-1].lower() in URL_TLDS:
return False # example.com/api is not a repo path
if head.upper() == head and len(head) > 2 and "." not in head:
return False # OPENAI_API_KEY is not a directory
...
return head.lower() in COMMON_TOP_DIRS or head in root_entries
Accept a token when it has a known extension, or when its first segment is a directory that
actually exists at the repo root, or a conventional one (src, docs, scripts, …). Placeholders
(<id>, {name}, path/to/…) are rejected — a harness that says path/to/secret.json is not
making a falsifiable claim.
Globs are claims too: src/**/*.tsx that matches nothing is drift, exactly as a missing file is.
2. command_runs — declared, never executed
This is the check people react to, because the naive version is a security incident: you do not
run the commands you are auditing.
if base in {"npm", "pnpm", "yarn", "bun"}:
...
if sub in {"run", "run-script", "rs"}:
script = tokens[2]
elif sub in NO_DECL_SUBCOMMANDS: # install, add, exec, doctor, ...
return CommandVerdict(True)
else:
script = sub # `pnpm build` == `pnpm run build`
if manifests.declares_script(script):
return CommandVerdict(True)
close = difflib.get_close_matches(script, sorted(manifests.scripts), n=1, cutoff=0.7)
return CommandVerdict(False, "high",
f"Script `{script}` is not declared in package.json", ...,
replace=(script, close[0]) if close else None)
Everything is a lookup against declarations the repo already makes: package.json scripts (across
workspaces), Makefile targets parsed with a line regex, just recipes, [project.scripts] entry
points from pyproject.toml via tomllib, and requirements*.txt / dependency tables for bare
binaries. A binary "resolves" if it is on PATH (augmented with node_modules/.bin, .venv/bin,
~/.local/bin, Homebrew dirs) or declared by a manifest — because a documented tool you simply
have not installed yet is not harness drift.
3. file_referenced — skills go both ways
Two independent failures hide in a skills folder:
- declared but unused: the harness names a skill, and the only mention of it anywhere in the repo is the harness line that declares it. Nothing outside the skill's own directory references it. That is a dead capability with a confident description.
-
on disk but unlisted:
.hermes/skills/orphan-skill/with a realSKILL.mdinside, and no instruction file that mentions it. The agent cannot use what it is never told about.
The first is why the exclusion predicate matters — you must exclude the skill's own files and the
declaration line, or the check passes vacuously on every skill.
4. server_reachable — schema plus binary, no spawning
.mcp.json, claude_desktop_config.json, .cursor/mcp.json, .vscode/mcp.json and opencode.json
all hold server maps under slightly different keys (mcpServers, servers, mcp). Validate the
shape — command is a non-empty string, args a list of strings, env an object — then resolve
the binary. Never start the server.
The diff for a dead server is the most satisfying piece of the codebase. Reparsing the JSON and
re-dumping it would reformat the whole file; instead a small brace-matching scanner finds the exact
line span of the member and deletes just that block, then re-parses the result and refuses to
emit the diff if the JSON would break:
patched = "\n".join(remaining) + "\n"
try:
json.loads(patched)
except ValueError:
return None
return make_diff(path, text, patched)
If the dead binary has a near neighbour on PATH (node21 → node), the patch rewrites the
command instead of deleting the server.
5. semantic — and unknown as a real answer
Conventions are the hard part, so they get two engines. A mechanical subset is checkable without a
model: never use X, use A instead of B, every unit must contain file.ext. Grep the
token through source files, skipping comments and docs — a README that quotes an anti-pattern to
tell you not to write it is not a violation.
Everything else goes to the model, with a strict contract:
{"verdicts":[{"id":"r7","verdict":"holds|violated|unknown","evidence":"<path:line + quote>"}]}
…and then the answer is audited too. If a violated verdict cites a path that does not exist in the
repo, it is downgraded to unknown and the unverified citation is kept on the finding. unknown
costs 1 point and is always displayed. A tool that hides its uncertainty is a tool that will
eventually hide a violation.
No key configured? The section reports skipped: no_api_key — as a row in the findings table, not a
silent gap. Checks 1–4 still run, which is where most real drift lives anyway.
Scoring you can defend
BASE_SCORE = 100
PENALTY_DETERMINISTIC_FAILURE = 8
PENALTY_SEMANTIC_VIOLATION = 5
PENALTY_SEMANTIC_UNKNOWN = 1
MAX_PENALTY_PER_CATEGORY = 32
CATEGORIES = ("paths", "commands", "skills", "mcp", "conventions")
score = max(0, 100 - sum(min(category_penalty, 32) for category in categories))
The per-category cap exists because a 500-path repo with one renamed directory should not score 0.
Four deterministic failures saturate a category, and the UI says capped when it happens.
The constants are served by GET /api/scoring and rendered in an expandable "how this score was
computed" panel, so nobody has to trust the number.
The two bugs that actually taught me something
lstrip("./") is not "strip a leading ./". It strips any leading character in the set
{'.', '/'} — so the pattern .mcp.json silently became mcp.json, and every dotted harness file
(.mcp.json, .cursorrules, .hermes/skills/**, .claude/…) was invisible to discovery. The
symptom was a clean-looking audit of a repo with a broken MCP config. The fix is a loop:
while norm.startswith("./"):
norm = norm[2:]
Worse than the MCP miss: inspect_repo(".eslintrc.json") returned "does not exist", so the auditor
would have confidently reported a false finding. A byte-cap or a regex bug is visible; a
path-normalisation bug manufactures evidence. That is the one class of bug this project cannot
afford.
A diff is a claim too. My first stale-path patch deleted the whole harness line. For
- Edit \src/old.tswhen the CLI output changes. that is right. For
Read \README.mdfirst, then \docs/architecture.md. it deletes a valid instruction along with
the stale one. So: if the line carries other checkable claims, excise only the stale token — which
means tidying the residue (… first, then . → … first.) with a small chain of regexes, and
falling back to deleting the line when nothing checkable survives.
And a rename is only proposed when the file name is unchanged (a move) or the two paths are ≥0.86
similar. Anything looser goes in the suggestion field, not in a diff. src/cli.ts → src/util.ts
is a plausible-looking patch that would be wrong, and a wrong patch is worse than no patch.
Precision: the audit of the audit
Luthier's first run against a real repo scored 42 and flagged src/old.ts — a path that repo's
README quotes as an example of drift. Technically a true observation; practically noise, and
noise in a score is how you lose people's trust in one edit.
Two rules fixed most of it:
-
A file the docs say the app writes is not a broken reference.
"stored in run/settings.local.json (mode 0600)"describes run-time behaviour. Advisory, 0 points. -
Prose docs naming a path inside a directory this repo does not have are probably examples.
README.mdquotingsrc/foo.tsin a repo with nosrc/at all → advisory. ButCLAUDE.mdquoting it → still a hard failure, because an instruction file is what the agent loads and obeys.
That distinction — instruction files are charged, docs are surfaced — is the whole difference
between a score people act on and a lint report nobody reads. The same repo went 42 → 76 with the
genuinely real findings still listed.
Verification
npm run verify # python -m agent.verify
Two throwaway git repos are written to a temp dir. One has five kinds of seeded drift; every one
must be flagged at the correct file:line. The other must score ≥95 with zero deterministic
failures — the guard against a checker that flags everything, which is easy to write and worthless.
Then every generated diff is run through git apply --check, every citation must be a real line in
a real file, and audit.log must be valid JSONL with tool, args, duration and bytes on every row.
drifted fixture score 47 findings 9 rules 9 tool calls 15
clean fixture score 100 findings 0 rules 4
30 passed, 0 failed
PASS
The Strands agent loop gets its own check (scripts/agent_loop_check.py) against a local mock
OpenAI-compatible endpoint, asserting the agent's tool call really travelled through the journaled
reader and that the API key never appears in a response body. It proves the integration, not the
model's judgement — and it is honest about that distinction, because a claim about the second kind
would be exactly the thing this tool exists to catch.
What it's for
Not lint. Lint checks that code satisfies a style. Luthier checks that instructions satisfy
reality — a different axis, and the one nobody automated yet.
The uncomfortable framing: a harness is a prompt, a prompt is a dependency, and we have no
dependency checker for prompts. We have this instead — a scorecard that says, in one number and a
table of quotes, how much of what you tell your agent is still true.
Run it against your repo. If the number is high, the interesting question is whether your harness
has been edited recently enough to have had the chance to rot.
Stack: Strands Agents 1.57 (agent loop, @tool), FastAPI + uvicorn, httpx for plain OpenAI-compatible chat calls, Vite + React 18 + TypeScript + Tailwind, difflib for the patches, git apply --check as the diff gate. Python 3.11+, no network from the checks, no writes to the
audited repo.
Code & more: https://www.dailybuild.xyz/project/271-luthier
Top comments (0)