DEV Community

Dekita
Dekita

Posted on Originally published at dekita-ai.pages.dev

Your coding agent is following rules that are no longer true. Here is the fix.

The most dangerous agent bug is not a model that invents an API. It is an instruction file that is still true enough to be trusted, and stale enough to be wrong.

Your AGENTS.md says the tests run with npm test. Last month a teammate moved them behind a build step. The agent reads the rule, trusts it, runs the old command, and reports a clean pass that never actually ran. No error is thrown. The stale rule is, by design, silent.

That is worse than having no rules at all. No rules make the agent conservative. A stale rule makes it confident.

The failure mode, concretely

An instruction file goes stale without failing. It references a path that moved, a command that changed, or a convention the team abandoned. Nothing checks whether the rule is still true, because the rule is plain prose with no owner and no schedule.

The result is a trust asymmetry: the file looks authoritative, so the agent follows it faithfully and confidently does the wrong thing. The maintainer has no signal that the guidance drifted, because the file does not throw on mismatch the way code does.

Treat the instruction file as code

The fix starts with a change of category. AGENTS.md is not documentation you write once and forget. It is a living input to a system that consumes it on every run. Treat it on the same maintenance cadence as your README or your API docs — which is to say, it needs a review trigger.

Two mechanisms keep it honest. One catches drift after the fact. The other verifies that a rule actually changes behavior before you trust it.

Mechanism one: a doc-drift check in CI

Give each agent-visible file a covers: line that lists the paths it describes. Then add a CI job that compares the last commit time of each doc against the last commit time of the paths it covers.

# scripts/check_doc_drift.sh
# fails when a doc is older than its subject by more than MAX_LAG_DAYS
MAX_LAG_DAYS=14
for doc in docs/agents/*.md; do
  covers=$(grep -oP '^covers:\s*\K.+' "$doc" || true)
  [ -z "$covers" ] && continue
  doc_ts=$(git log -1 --format=%ct -- "$doc")
  for target in $covers; do
    target_ts=$(git log -1 --format=%ct -- "$target")
    lag=$(( (target_ts - doc_ts) / 86400 ))
    if [ "$lag" -gt "$MAX_LAG_DAYS" ]; then
      echo "STALE: $doc is $lag days behind $target"; exit 1
    fi
  done
done
echo "Docs OK"
Enter fullscreen mode Exit fullscreen mode

The one gotcha that will bite you: the git checkout in CI must use fetch-depth: 0. Otherwise git log sees a single commit since the shallow clone, every doc looks freshly updated, and the job always passes while the docs quietly rot.

Mechanism two: verify a rule by discarding and retrying

A doc-drift check catches a file that fell behind. It does not tell you whether a new rule does anything. The only way to learn that is to test the rule in isolation.

1. Add or modify the rule
2. Discard the current artifact (or stash it on a branch)
3. Start a fresh session with the updated rules
4. Re-run the same task
5. Confirm the issue does not recur
Enter fullscreen mode Exit fullscreen mode

If you keep the existing artifact and continue, you are still operating in context polluted by the old system. The model may try to reconcile the new rule with the old work instead of applying it cleanly. You cannot tell whether the rule works, or whether you just fixed the symptom by hand. Discard-and-retry is the only clean measurement.

What belongs in the file at all

Before you add a rule, ask whether the problem is best solved by a line of prose. If a rule can be expressed as a test, a hook, or a permission boundary, write it there instead — those fail loudly when they rot. Reserve the prose for what only prose can carry:

  • An expensive operation the agent must never run by accident
  • Code the agent must not touch
  • A project-level security boundary
  • A convention that lives only in the team's heads and is not visible in the code

Everything else is context the agent can read from the code itself. A rule that restates the code adds tokens without adding signal.

Add a verification checklist to the file

An instruction file should end the way a pull request does: with a confirmation step. A short checklist forces the agent to verify before it reports done, and it gives a human something concrete to review in the diff.

## Before you finish
- [ ] The command in step 4 matches the one in CI
- [ ] The path in step 2 still exists on the main branch
- [ ] You can point to the test that would fail if this rule broke
Enter fullscreen mode Exit fullscreen mode

When the file ships with a check, a stale rule stops being silent. It becomes a checkbox nobody can honestly tick, which is exactly the signal a maintainer needs.

A stale rule in the wild

Here is the shape of the failure, the one that made me stop trusting my own instructions file. The repo had a rule: "Run make check before you finish." A teammate reorganized the Makefile and renamed the target to make ci-check. Nothing referenced the old name, so nothing failed. The agent read the rule, ran make check, got a target-not-found error — and, because the rule said "before you finish," decided the error was a pre-existing environment issue and reported done anyway. It did not update the rule, because the rule did not tell it that it was allowed to.

That is the silent-drift failure in one loop. The rule was still on the page, still trusted, and wrong. No test, no linter, no human was in the path to catch it. It was not a discipline failure on anyone's part. The system had no place where staleness was supposed to be detected.

The maintenance cadence that makes it honest

A drift check is only useful if it runs. Put it on the same schedule as your dependency updates and your license scans, not as a one-time cleanup you do when you remember.

- Every PR that touches a path a doc covers → re-run the drift check
- Every model-version upgrade → re-read AGENTS.md and delete rules written for the old model
- Every month → a rule you cannot remember triggering is a rule you delete
Enter fullscreen mode Exit fullscreen mode

The last one is the hard one. Most rules are added reactively, after a bug. The bug stops happening, the rule stays, and a year later it is a small tax on every session. A rule with no documented rationale and no recent trigger is not insurance, it is rent. Delete it, and let the next real bug re-earn its place.

Why this is a design problem, not a discipline problem

The instinct is to blame the team: someone should have updated the file. But the real issue is that the file had no mechanism to announce its own staleness. A doc with no owner, no covers: line, and no checklist is structurally unable to tell you when it is wrong. You cannot discipline your way around a missing feedback loop.

Where I am unsure

I have not measured the failure rate of stale rules on a large corpus. The covers:-line drift check and the discard-and-retry loop are practices I have seen work in small teams, not numbers I can cite. Treat the mechanism as the argument, and the specific thresholds as a starting point you tune to your repo.

What I am confident about is the structure: an instruction file that is trusted but not verified will fail in the direction of the old world. The fix is to give the file the same verification discipline you give the code — a review trigger, and a way to prove a rule still works.

Sources

  • dev.to, "AGENTS.md Pitfalls: 7 Mistakes That Make Coding Agents Less Reliable" (2026)
  • dev.to, "How to write an AGENTS.md your AI agent actually follows" (2026)
  • dev.to, "Agents Don't Need Memory, They Need Documentation: A Practical AGENTS.md Playbook" (2026)
  • dev.to, "Stop Putting Everything in AGENTS.md" (2026)

This post was written with AI assistance. The author is responsible for its content.

Top comments (1)

Some comments may only be visible to logged-in visitors. Sign in to view all comments.