DEV Community

Cover image for What Actually Belongs in AGENTS.md: My Production Next.js + TypeScript Setup (and the CI Check That Keeps It Honest)
Nael M. Awadallah
Nael M. Awadallah

Posted on

What Actually Belongs in AGENTS.md: My Production Next.js + TypeScript Setup (and the CI Check That Keeps It Honest)

A complete, annotated AGENTS.md for production Next.js + TypeScript: what I keep, what I cut, how it behaves across Claude Code, Codex, and Cursor, and a CI script that fails the build when it drifts.

In my previous article, I explained how I separate shared instructions, tool-specific configuration, intent documents, and skills across Claude Code, Codex, and Cursor.

The most common follow-up question was simple:

"OK, but what actually goes in the AGENTS.md?"

This article answers that with a complete file, not fragments. You'll get:

  1. The one test I apply to every line
  2. The full AGENTS.md for a production Next.js + TypeScript app
  3. Why each section exists
  4. What I deliberately leave out, and where it goes instead
  5. How nested AGENTS.md files behave in Claude Code, Codex, and Cursor (they differ)
  6. A zero-dependency CI script that fails the build when your AGENTS.md lies

Where this comes from: The setup in this article comes from a production client project: a Next.js App Router frontend with TypeScript in strict mode, built mostly with Claude Code. The project is under NDA, so the example below is generalized and names are changed.

Section 2 is a template you can adapt to any Next.js + TypeScript project. The example uses pnpm, Zod, Vitest, and Playwright; swap in your own tools.

Tool behavior was checked against official documentation in September 2026. Claude Code reads AGENTS.md directly only from v2.1.277 onward.


1. The Only Test That Matters

Before any line goes into AGENTS.md, I ask:

"Would the agent get this wrong without this line?"

If the answer is no, the line is noise. It costs context on every task and dilutes the lines that matter.

That single question eliminates most of what people put in these files:

  • "Write clean, maintainable code." The agent was going to try anyway.
  • "Use TypeScript." It can see tsconfig.json.
  • A folder tree. It can list the directory.

And it keeps what agents genuinely can't infer:

  • Which of three plausible test commands is the right one
  • That src/server/ must never be imported by client code
  • That Server Actions need authorization checks even when the UI hides the button

Anthropic's Claude Code documentation suggests a practical trigger for adding instructions: when the agent makes the same mistake a second time. I use that as my "two strikes" rule. One mistake might be noise. Two is a missing instruction.

A real example: four corrections, four different homes

On a recent client project, I asked Claude Code to plan a set of dashboard screens. The plan was architecturally sound, but review surfaced four corrections. They turned out to be the best illustration of this whole article, because each one belonged in a different place:

Correction Where it belongs
The plan added a new test framework as the last item of a task list. A new tool deserves an explicit decision, not a line at the bottom of a plan. AGENTS.md: "Present new dependencies or tooling as a separate decision requiring approval."
Mock data was shaped as ready-to-render view models. When the real endpoints arrived, every fixture would need rewriting. AGENTS.md: "Shape fixtures like raw API responses and pass them through the mapping layer."
Shared widgets were planned as module-local components, even though several screens needed them. Architecture docs: component placement rules
A panel appeared in the design without a confirmed product need. Not an instruction at all: a product question, resolved in the feature's intent document

Only two of the four became permanent rules. If I had put all four into AGENTS.md, I'd have mixed repository conventions with one feature's decisions, which is exactly the problem from my previous article.


2. The Complete File

Here's the full AGENTS.md. It's under 100 lines.

# Acme Dashboard

Next.js (App Router) + TypeScript (strict). pnpm only.
Check the `next` version in package.json before using
version-specific APIs. Do not rely on memory for them.

## Commands

- Install: pnpm install
- Dev server: pnpm dev
- Typecheck: pnpm typecheck
- Lint: pnpm lint
- Unit tests (all): pnpm test
- Unit tests (one file): pnpm test <path-to-test-file>
- E2E: pnpm test:e2e (requires the dev server running)

## Where things live

- Server-only code (DB access, secrets, auth checks): src/server/
- Zod schemas: src/lib/schemas/
- Shared UI primitives: src/components/ui/ (reuse before creating)
- External backend client: src/lib/api/client.ts

## Next.js rules

- Default to Server Components. Add "use client" only for
  state, effects, event handlers, or browser APIs, and push
  it down to the smallest leaf component.
- Every file in src/server/ must `import "server-only"`.
- Only NEXT_PUBLIC_ variables reach the browser.
  Never give a secret that prefix.
- Server Actions and Route Handlers are public HTTP endpoints.
  Authenticate and authorize inside every one, even when the
  UI hides the trigger.
- Do not change caching or revalidation behavior unless the
  task asks for it. Defaults differ between Next.js versions.

## TypeScript rules

- No `any`. No `@ts-ignore`. If a suppression is unavoidable,
  use `@ts-expect-error` with a one-line reason.
- Do not use `as` to silence errors. Fix or narrow the type.
- Parse all external input (request bodies, search params,
  API responses) with Zod and infer types with z.infer.

## Boundaries

Ask before:
- Adding, removing, or upgrading dependencies or tooling
  (present it as a separate decision, not a task-list item)
- Creating or editing database migrations
- Changing authentication, permissions, or middleware
- Editing .env files, CI configuration, or the lockfile by hand

Never:
- Commit secrets or real customer data, including in fixtures
- Disable lint rules or skip tests to make a check pass

No need to ask:
- Run typecheck, lint, and unit tests, and fix failures
  caused by your own change

## Mock data

- Shape fixtures like raw API responses and pass them
  through the mapping layer, never as pre-mapped view models.

## Definition of done

- pnpm typecheck and pnpm lint pass
- Tests for affected modules pass; new logic has tests
- Report anything you could not verify

## Docs (read when relevant, not before every task)

- docs/architecture.md: module boundaries and data flow
- docs/auth.md: sessions, roles, permission checks
- docs/api.md: external backend contract and error format
- intent/: requirements for a feature, when a task references one

## Code Review Rules

- Flag Server Actions or Route Handlers without an
  authorization check.
- Flag client components that import from src/server/.
- Flag new NEXT_PUBLIC_ variables that look like secrets.
Enter fullscreen mode Exit fullscreen mode

Now let's go through why each section earns its place.


3. Section by Section

The header: two lines, one of them a warning

Next.js (App Router) + TypeScript (strict). pnpm only.
Check the `next` version in package.json before using
version-specific APIs. Do not rely on memory for them.
Enter fullscreen mode Exit fullscreen mode

"pnpm only" prevents a surprisingly common failure: an agent running npm install and generating a second lockfile.

The version line matters more than it looks. Next.js has changed important defaults between major versions. A well-known example: in Next.js 15, params, searchParams, cookies(), and headers() became asynchronous. Caching defaults have also changed across versions 14, 15, and 16.

Models are trained on years of code written for older versions, so they confidently produce older patterns.

That's why the file doesn't name a Next.js version or list what changed. Both would go stale on the next upgrade, and a list written from memory could itself be wrong. Instead, the file tells the agent where the truth lives: the installed version in package.json, and the official docs for that version.

Commands: the one thing agents truly can't guess

- Unit tests (one file): pnpm test <path-to-test-file>
- E2E: pnpm test:e2e (requires the dev server running)
Enter fullscreen mode Exit fullscreen mode

Agents can guess pnpm test. They can't guess:

  • How to run a single test file, which is what they need most during iteration
  • That E2E tests need a running server
  • That typecheck is a separate script from lint

Note the <path-to-test-file> placeholder. I use angle brackets for placeholders so they're obviously not real paths. This matters for the CI check in section 7.

Where things live: only the non-obvious

I don't include a folder tree. The agent can run ls.

I include locations that carry rules: where server-only code lives, where schemas live, which UI primitives to reuse. Each line either prevents a duplicate or protects a boundary.

Next.js rules: security first, style never

This section is short on purpose, and every line maps to a real failure mode:

Rule Failure it prevents
Default to Server Components, push "use client" to leaves Entire pages becoming client bundles
import "server-only" in src/server/ Database or secret code leaking into a client bundle. The import makes this a build error instead of a silent leak.
NEXT_PUBLIC_ warning Secrets inlined into JavaScript shipped to every browser
Server Actions are public endpoints Authorization enforced only in the UI, bypassable with a direct POST request
Don't touch caching unless asked "Fixes" that change data freshness across the app

The Server Actions rule is the one I'd keep if I could keep only one. An agent asked to "add a delete button for admins" will naturally hide the button for non-admins and consider the job done. But the action itself remains callable by anyone who can send a request.

Hiding actions based on permissions is still good practice; users shouldn't see buttons they can't use. But that's user experience, not security. The server has to enforce authorization on every request, whether it's a Server Action, a Route Handler, or a separate backend API.

So the rule I give agents is short:

Hiding a button is user experience. Authorization happens on the server, every time.

TypeScript rules: close the escape hatches

Agents under pressure to make tsc pass reach for three shortcuts: any, @ts-ignore, and as. All three make the error disappear without fixing anything.

Requiring @ts-expect-error with a reason, instead of @ts-ignore, has a useful property: if the underlying error is ever fixed, TypeScript reports the directive as unused, so stale suppressions don't accumulate silently.

Boundaries: three tiers, not a wall of NEVER

Many instruction files look like this:

NEVER modify the database.
NEVER change dependencies.
ALWAYS ask before doing anything significant.
Enter fullscreen mode Exit fullscreen mode

This creates two problems. The agent can't distinguish "irreversible and dangerous" from "mildly annoying," and newer models can take heavy-handed restrictions so literally that they stop and ask about routine work. OpenAI's guidance for recent models explicitly recommends revisiting overly strong boundary language for this reason.

So I use three tiers:

  • Ask before: consequential but legitimate changes (dependencies, migrations, auth, lockfile)
  • Never: actions with no valid reason in normal work
  • No need to ask: explicit permission for safe, repetitive work, so the agent doesn't stop after every test run

The third tier is the one most people forget. Granting permission is as important as restricting it.

Notice the detail in the first tier: new tooling should be presented "as a separate decision, not a task-list item." That line comes directly from the plan review in section 1. A dependency buried at the bottom of a ten-step plan tends to get approved along with everything else, without anyone deciding on it.

Mock data: one line that prevents a rewrite

The fixtures rule is the other line from that plan review. Mock data shaped like the final UI feels convenient, but when the real API arrives, every fixture and every consumer needs rewriting. Mock data shaped like real API responses, passed through the same mapping layer the live data will use, lets you swap the data source at one point without touching the rest of the feature.

Definition of done

Without a definition of done, "done" means "the agent stopped." Three lines fix that, and the last one ("report anything you could not verify") is the most valuable. It turns silent gaps into visible ones.

Docs: a router, not a reading list

- docs/auth.md: sessions, roles, permission checks
Enter fullscreen mode Exit fullscreen mode

Each entry says when the document is relevant. I never write "read these files before every task." That burns context on typo fixes, and capable models are good at pulling the right document when they know what each one covers.

Code Review Rules

Codex's GitHub code review reads a ## Code Review Rules section from the AGENTS.md closest to the changed code. Even if you don't use Codex review, this section is useful: it tells any agent what reviewers will flag, so it can avoid those issues in the first place.

Keep these rules about behavior that CI can't catch. Formatting belongs to your linter.


4. What I Deliberately Leave Out

This is the most important section of the article. Every item below is something I've seen in real AGENTS.md files.

Left out Why Where it goes instead
Formatting and style rules Linters and formatters enforce them deterministically. An instruction is only a suggestion. ESLint, Prettier, CI
Folder trees Stale after the next refactor; the agent can list directories Nowhere
Generic advice ("write clean code") Changes nothing; costs context Nowhere
Full architecture explanations Too long for every-task context docs/, linked from the router section
Feature requirements Temporary; pollute permanent instructions intent/ documents
Multi-step procedures (release, migration review) Needed occasionally, not on every task Skills
Personal preferences Not shared team knowledge CLAUDE.local.md, user-level config
Secrets, URLs, credentials Committed to the repository Environment variables, secret manager
Hardcoded version numbers Go stale silently "Check package.json"
@path imports Claude Code syntax; other agents see plain text, and Claude loads the file at launch, defeating on-demand reading Plain-text doc references
"Read X before every edit" Burns context on every trivial task "Use X when doing Y"
Framework tutorials The model already knows the framework; it needs your deviations from it Nowhere

Two items deserve a closer look.

@path imports. In Claude Code, @docs/auth.md inside an instruction file means "load this file into context at launch." That's useful in CLAUDE.md, but in a shared AGENTS.md it causes two problems. It loads the document on every Claude session, even when irrelevant. And Codex or Cursor may treat it as plain text, so the instruction means different things to different agents. In a shared file, reference documents in plain text and let each agent read them when needed.

Linter-enforceable rules. If a rule can be enforced by a tool, enforce it with the tool. An agent following an instruction is probabilistic. A failing lint check is not. And once the linter enforces it, the agent learns the rule from the error message anyway.


5. Nested AGENTS.md Files: Same File, Three Behaviors

In larger apps or monorepos, you might add AGENTS.md files in subdirectories:

repo/
├── AGENTS.md                 # repo-wide rules
├── apps/
│   ├── web/
│   │   └── AGENTS.md         # Next.js app rules
│   └── admin/
│       └── AGENTS.md         # admin app rules
└── packages/
    └── ui/
        └── AGENTS.md         # component library rules
Enter fullscreen mode Exit fullscreen mode

Here's the part most guides miss: the three tools load these differently. The table below reflects each tool's official documentation as of September 2026. This is the area most likely to change, so re-check it after major upgrades.

Tool When is apps/web/AGENTS.md used?
Codex Only when you start Codex in apps/web/ or below. Codex walks from the repo root down to the launch directory, one file per directory. Launch from the root, and it won't include apps/web/AGENTS.md.
Cursor When the agent works with files in apps/web/. Nested files combine with parent files; more specific instructions take precedence.
Claude Code (reading AGENTS.md directly) When Claude reads a file in apps/web/ and that directory has no CLAUDE.md of its own.
Claude Code (with a root CLAUDE.md) By default, not read. A root CLAUDE.md switches Claude to reading CLAUDE.md files only.

That last row is a trap if you followed my previous article's recommendation to add a root CLAUDE.md with @AGENTS.md. The import covers the root file, but nested AGENTS.md files stop being read.

Two fixes:

  1. Add a small CLAUDE.md next to each nested AGENTS.md, containing just @AGENTS.md. Claude loads subdirectory CLAUDE.md files on demand when it works there.
  2. Change Claude Code's default. Run /config and set Project instructions to claude-md-and-agents-md, which loads both kinds of files.

My rules for nested files:

  • Never contradict the parent. Nested files should narrow or add, not override. Contradictions produce behavior that depends on which tool you're using.
  • Keep them self-contained. Because Codex only sees them from the right launch directory, don't let critical rules exist only in a nested file if people routinely launch from the root.
  • Watch the total size. Codex concatenates the chain and truncates it at 32 KiB by default (configurable with project_doc_max_bytes). Content past the limit is silently dropped.

6. Using It Across Agents

With the file in place, here's the minimal wiring per tool.

Claude Code: root CLAUDE.md:

@AGENTS.md

## Claude Code

- For changes under src/server/auth/, propose a plan before editing.
Enter fullscreen mode Exit fullscreen mode

Verify with /memory or /context.

Codex: no wiring needed. Launch Codex from the directory whose rules should apply, and verify with:

codex --ask-for-approval never "List the instruction sources you loaded."
Enter fullscreen mode Exit fullscreen mode

Cursor: reads AGENTS.md automatically. I add .cursor/rules/*.mdc files only for glob-scoped rules that shouldn't load on every task. Check active rules under Customize → Rules.

The principle from the previous article holds: one shared file, and tool-specific files only for tool-specific behavior.


7. The Surprise: A CI Check That Keeps AGENTS.md Honest

Here's the problem with instruction files: they rot silently.

Someone renames docs/auth.md to docs/authentication.md. Someone renames the typecheck script to check:types. Nobody updates AGENTS.md, because nothing fails. The agent now gets instructions pointing to files that don't exist and commands that error out, and it has no way to know the instructions are wrong.

Code has tests. Instruction files should too.

This script, scripts/check-agents-md.mjs, has zero dependencies and runs on Node 18+. It checks every AGENTS.md in the repository for:

Check Severity
Referenced paths (docs/…, src/…, intent/…) that don't exist Error
pnpm/npm/yarn/bun run commands with no matching package.json script Error
Instruction chain larger than Codex's 32 KiB default limit Error
Files over 200 lines Warning
@path imports in shared files Warning
Root CLAUDE.md that doesn't import @AGENTS.md Warning
CLAUDE.local.md present (silently disables AGENTS.md loading in Claude Code) Warning
AGENTS.override.md present (Codex ignores the sibling AGENTS.md) Warning
#!/usr/bin/env node
// scripts/check-agents-md.mjs
// Validates AGENTS.md files: size budget, broken path references,
// missing package scripts, and cross-tool loading pitfalls.
// Zero dependencies. Node 18+.

import { readFileSync, readdirSync, existsSync, statSync } from "node:fs";
import { join, dirname, relative, resolve } from "node:path";

const ROOT = process.cwd();
const CODEX_DEFAULT_MAX_BYTES = 32 * 1024; // Codex project_doc_max_bytes default
const WARN_LINES = 200;
const IGNORED_DIRS = new Set([
  "node_modules", ".git", ".next", "dist", "build", "coverage", ".turbo", ".vercel",
]);
const INSTRUCTION_FILES = new Set(["AGENTS.md", "AGENTS.override.md"]);
const PM_BUILTINS = new Set([
  "install", "i", "ci", "add", "remove", "rm", "exec", "dlx", "update", "up",
  "why", "list", "ls", "audit", "outdated", "create", "init", "link", "unlink",
  "publish", "store", "config", "run", "x",
]);
const PATH_RE =
  /(?:^|[\s`'"(])((?:\.\/)?(?:docs|src|app|lib|intent|scripts|packages|apps|tests?|e2e)\/[^\s`'"]*)/gm;
const SCRIPT_RE = /\b(npm|pnpm|yarn|bun)\s+(run\s+)?([a-z][\w:.-]*)/gi;
const IMPORT_RE = /(?:^|\s)@([\w./-]+\.\w+)/gm;

const errors = [];
const warnings = [];
const rel = (p) => relative(ROOT, p) || ".";

function findFiles(dir, out = []) {
  for (const entry of readdirSync(dir, { withFileTypes: true })) {
    if (entry.isDirectory()) {
      if (!IGNORED_DIRS.has(entry.name)) findFiles(join(dir, entry.name), out);
    } else if (entry.isFile() && INSTRUCTION_FILES.has(entry.name)) {
      out.push(join(dir, entry.name));
    }
  }
  return out;
}

function nearestPackageJson(fromDir) {
  let dir = fromDir;
  while (true) {
    const candidate = join(dir, "package.json");
    if (existsSync(candidate)) return candidate;
    if (dir === ROOT || dir === dirname(dir)) return null;
    dir = dirname(dir);
  }
}

function readScripts(pkgPath) {
  try {
    return JSON.parse(readFileSync(pkgPath, "utf8")).scripts ?? {};
  } catch {
    errors.push(`${rel(pkgPath)}: could not parse package.json`);
    return null;
  }
}

// Approximates what Codex concatenates when started in `dir`:
// one file per directory from the repo root down (override wins).
// Does not include your global ~/.codex/AGENTS.md.
function codexChainBytes(dir) {
  let total = 0;
  let current = dir;
  while (true) {
    const override = join(current, "AGENTS.override.md");
    const base = join(current, "AGENTS.md");
    if (existsSync(override) && statSync(override).size > 0) total += statSync(override).size;
    else if (existsSync(base)) total += statSync(base).size;
    if (current === ROOT || current === dirname(current)) break;
    current = dirname(current);
  }
  return total;
}

function cleanPath(raw) {
  let p = raw.replace(/[.,;:!?]+$/, "");
  const count = (s, ch) => s.split(ch).length - 1;
  while (p.endsWith(")") && count(p, ")") > count(p, "(")) p = p.slice(0, -1);
  return p.replace(/[.,;:!?]+$/, "");
}

function checkFile(file) {
  const dir = dirname(file);
  const text = readFileSync(file, "utf8");
  const name = rel(file);
  const lines = text.split("\n").length;

  if (lines > WARN_LINES) {
    warnings.push(`${name}: ${lines} lines (aim for under ${WARN_LINES}; longer files reduce adherence)`);
  }

  const chain = codexChainBytes(dir);
  if (chain > CODEX_DEFAULT_MAX_BYTES) {
    errors.push(
      `${name}: instruction chain is ${chain} bytes; Codex truncates at ${CODEX_DEFAULT_MAX_BYTES} by default`,
    );
  }

  for (const [, raw] of text.matchAll(PATH_RE)) {
    const p = cleanPath(raw);
    if (!p || /[*{}<>]/.test(p)) continue; // skip globs and placeholders
    if (!existsSync(resolve(ROOT, p)) && !existsSync(resolve(dir, p))) {
      errors.push(`${name}: references missing path "${p}"`);
    }
  }

  const pkgPath = nearestPackageJson(dir);
  const scripts = pkgPath ? readScripts(pkgPath) : null;
  if (scripts) {
    for (const [, pm, runKeyword, script] of text.matchAll(SCRIPT_RE)) {
      const manager = pm.toLowerCase();
      const hasRun = Boolean(runKeyword);
      if (!hasRun && PM_BUILTINS.has(script)) continue;
      if (manager === "npm" && !hasRun && !["test", "start"].includes(script)) continue;
      if (manager === "bun" && !hasRun) continue;
      if (!(script in scripts)) {
        errors.push(`${name}: "${pm} ${hasRun ? "run " : ""}${script}" has no matching script in ${rel(pkgPath)}`);
      }
    }
  }

  for (const [, target] of text.matchAll(IMPORT_RE)) {
    warnings.push(
      `${name}: "@${target}" is Claude Code import syntax; other agents read it as plain text, and Claude loads it at launch`,
    );
  }

  if (file.endsWith("AGENTS.override.md")) {
    warnings.push(`${name}: override file present; Codex ignores the AGENTS.md in this directory`);
  }
}

function checkClaudeWiring() {
  if (!existsSync(join(ROOT, "AGENTS.md"))) return;
  for (const candidate of ["CLAUDE.md", join(".claude", "CLAUDE.md")]) {
    const p = join(ROOT, candidate);
    if (existsSync(p) && !/(^|\s)@AGENTS\.md\b/m.test(readFileSync(p, "utf8"))) {
      warnings.push(
        `${candidate}: exists but does not import @AGENTS.md; by default Claude Code will not read AGENTS.md`,
      );
    }
  }
  if (existsSync(join(ROOT, "CLAUDE.local.md"))) {
    warnings.push(
      "CLAUDE.local.md: present; unless CLAUDE.md imports @AGENTS.md, Claude Code stops reading AGENTS.md for you",
    );
  }
}

const files = findFiles(ROOT);
if (files.length === 0) {
  console.error("error No AGENTS.md files found");
  process.exit(1);
}

files.forEach(checkFile);
checkClaudeWiring();

for (const w of warnings) console.warn(`warn  ${w}`);
for (const e of errors) console.error(`error ${e}`);
console.log(`\nChecked ${files.length} file(s): ${errors.length} error(s), ${warnings.length} warning(s)`);
process.exitCode = errors.length > 0 ? 1 : 0;
Enter fullscreen mode Exit fullscreen mode

Add it to package.json:

{
  "scripts": {
    "check:agents": "node scripts/check-agents-md.mjs"
  }
}
Enter fullscreen mode Exit fullscreen mode

And run it on every pull request:

# .github/workflows/agents-md.yml
name: agents-md
on: [pull_request]

jobs:
  check:
    runs-on: ubuntu-latest
    steps:
      - uses: actions/checkout@v4
      - uses: actions/setup-node@v4
        with:
          node-version: 20
      - run: node scripts/check-agents-md.mjs
Enter fullscreen mode Exit fullscreen mode

It runs on every PR, not only when AGENTS.md changes, because the most common breakage is a rename elsewhere that invalidates a reference.

Example output on a broken file:

warn  CLAUDE.md: exists but does not import @AGENTS.md; by default Claude Code will not read AGENTS.md
error AGENTS.md: references missing path "docs/auth.md"
error AGENTS.md: "pnpm typecheck" has no matching script in package.json

Checked 2 file(s): 2 error(s), 1 warning(s)
Enter fullscreen mode Exit fullscreen mode

Known limitations, so you can adapt it:

  • Path detection covers common top-level folders (docs/, src/, app/, intent/, and so on). Add your own to PATH_RE.
  • Paths containing *, {}, or <> are skipped as globs or placeholders. That's why the AGENTS.md above writes <path-to-test-file>.
  • The size check approximates Codex's chain and excludes your global ~/.codex/AGENTS.md.
  • CLAUDE.local.md is usually gitignored, so that warning only fires locally. Run the script locally too.

8. The Maintenance Loop

An AGENTS.md is never finished. Mine evolves through four habits:

  1. Two strikes, then a rule. When an agent repeats a mistake, I add the smallest line that would have prevented it.
  2. Every line needs an owner. I add AGENTS.md to CODEOWNERS. Instructions that shape every agent's behavior deserve the same review as code.
  3. Re-audit on model upgrades. Instructions written to compensate for an older model's weaknesses can over-constrain a newer one. When I switch models, I re-run the test from section 1 on every line.
  4. Let CI catch the rot. The script from section 7 handles broken references, so reviews can focus on whether instructions are still right, not whether they still exist.

9. Checklist

Before committing an AGENTS.md:

  • [ ] Every line passes: "Would the agent get this wrong without it?"
  • [ ] Commands include how to run a single test
  • [ ] Boundaries have three tiers: ask before, never, no need to ask
  • [ ] New dependencies and tooling require an explicit, separate decision
  • [ ] Security rules cover the server/client boundary and endpoint authorization
  • [ ] A definition of done exists
  • [ ] Docs are referenced by when to use them, not "read before every task"
  • [ ] No style rules a linter could enforce
  • [ ] No @path imports in the shared file
  • [ ] No hardcoded framework versions
  • [ ] Nested files narrow the parent, never contradict it
  • [ ] Claude Code wiring verified (/memory)
  • [ ] check-agents-md passes in CI

Final Thoughts

A good AGENTS.md isn't a description of your project. It's a list of the places where a capable engineer who has never seen your codebase would get it wrong.

That framing changes what you write. You stop explaining Next.js and start explaining your deviations from it. You stop listing folders and start protecting boundaries. You stop writing NEVER in capital letters and start granting permission for safe work.

And because the file shapes every agent's behavior on every task, it deserves what we give any other critical code: review, ownership, and a test that fails when it breaks.


Further Reading


Your Turn

What's the one line in your AGENTS.md that has saved you the most trouble?

And what did you remove that you thought you needed?

Share it in the comments. I'm collecting the best examples for a follow-up.

Top comments (0)