DEV Community

Cover image for Claude Code Review in Production: How Our Review Rules Evolved Over Five Months
Rohan Sen Sharma
Rohan Sen Sharma

Posted on Originally published at nulltensor.com

Claude Code Review in Production: How Our Review Rules Evolved Over Five Months

Claude Code review started at my company as a single prompt in a GitHub workflow, and five months later it is a playbook, six specialist agents and a verifier that throws out any finding it cannot prove. Almost every step between those two points was a reaction to something that went wrong: false positives while the rules were still settling, findings on code the pull request never touched, severity that depended on the model's mood, and a reviewer that invented problems when we gave it too many rules at once.

This is how the rules evolved and what each change fixed.

Claude Code's built-in review, and why we wrote our own rules

Claude Code now has a review of its own. Anthropic's managed Code Review, in research preview for Team and Enterprise plans, runs through the Claude GitHub App: several agents look for different classes of issue in parallel, a verification step filters out false positives, and findings arrive as inline comments tagged Important, Nit or Pre-existing. It never approves or blocks a pull request, and you tune it with a REVIEW.md file. Locally, /code-review does the same for your branch.

We kept our own because we already had a customised flow. The built-in review hunts for correctness bugs by default; ours enforces a written architecture, cites a rule for every finding, and votes on the pull request. If you want a second pair of eyes for bugs, start there; if you want documented rules enforced with a verdict, you need a playbook of your own.

Version 1: one prompt, run on request (April)

The first version went in on 20 April. A workflow listened for a pull request comment containing "@claude" and "review", and ran Claude Code with one prompt that did three things.

  • For files in our layered modules (routers, services, repositories and DTOs), it read the architecture rule documents for each layer and treated any violation as blocking.
  • Everything else got a general review, explicitly without the layering rules.
  • It grouped findings by file, each with a severity (BLOCKING, WARNING or SUGGESTION), a location, the problem and a concrete fix.

That structure survived every later version; the trigger did not. While the rules were settling it produced false positives, so it ran only when someone asked.

Version 2: every pull request, with a verdict (May)

Between 12 and 15 May the workflow changed four times.

  • It ran on every pull request, on open, on each new push and on reopen.
  • It submitted a real GitHub review: request changes if there was any blocking finding, approve otherwise.
  • It skipped draft pull requests until they were marked ready for review.
  • It skipped long-lived project branches, where every intermediate push would have drawn a review.

The lesson of that week: once a reviewer votes, noise stops being a nuisance and becomes a cost. A wrong "request changes" blocks a colleague. Everything that followed was about precision.

Version 3: load only the rules that matter, and review only what changed

The next rewrite introduced three ideas that are still the backbone of the system.

Load rules by path, not all at once. The reviewer lists the changed files, then reads only the rule documents those paths need: router rules only if a router changed, and so on. Loading every rule at once caused hallucinations: the reviewer reported violations that were not there.

Review only changed lines. The prompt gained a sentence we have never removed: review only changed or added lines, never flag pre-existing code. Without it, a one-line fix in an old file could come back with twenty findings about code written years earlier.

Severity by definition, not judgement. The severity section was headed "apply exactly, no judgment calls". BLOCKING meant a layering-rule violation, a correctness bug, a security issue or a data-loss risk. WARNING meant correct architecture that was still risky. SUGGESTION meant style or naming that broke no rule.

Version 4: one playbook for local and CI reviews

On 18 May the rules moved out of the workflow and out of CLAUDE.md into a single docs/review.md. Both the GitHub Action and a local /review command follow the same file, so a local review before pushing applies the same rules the pull request will face.

  • Ask for the base branch. Locally, the review asks which branch the work will merge into instead of assuming main.
  • Skip data files. SQL, CSV and JSON files are excluded; reviewing migrations and fixtures as code produced findings nobody could act on.
  • Every finding cites a rule. BLOCKING must cite an architecture rule file and section, WARNING a style rule from CLAUDE.md. SUGGESTION is for anything no rule covers.
  • Never block legacy code on style. Style caps at WARNING everywhere.

Tying severity to a citation changed the character of the reviews. A finding either points at a rule the team agreed, or it is a suggestion the author can ignore.

Version 5: specialists and a verifier

On 19 May the single reviewer became an orchestrator that reviews nothing itself. It routes each changed file to the specialists that own it, runs them in parallel and assembles the result.

Specialist Reviews Owns
Router reviewer API and router files in the layered modules Router-layer rules
Service reviewer Service files in the layered modules Service-layer rules
Repository reviewer Repository files in the layered modules Repository-layer rules
Architecture reviewer Any file in the layered modules Cross-cutting architecture rules
Python style reviewer Every Python file Style rules from CLAUDE.md, capped at WARNING
Verifier Every finding Nothing: it only checks the others' work

Each specialist reads only its own rule document, so a service file reviewed by four specialists does not get the same finding four times. They all return findings in one fixed format, with the offending code quoted verbatim.

The verifier is the Claude Code review agent I would copy first. For every finding it checks four things, and rejects the finding if any fails:

  1. The cited file exists.
  2. The quoted code appears within five lines of the cited line.
  3. That line was added or changed in this diff.
  4. The cited rule section exists and says what the finding claims.

Its instructions open with the principle behind the whole design: false positives are far worse than false negatives. It may not re-grade severity or add findings of its own, and rejected findings are dropped silently.

Adding a specialist without touching the workflow (September)

In September we added a sixth reviewer for database performance, because our most expensive production problems are queries that are fast in development and slow at scale. It blocks queries inside loops, N+1 access to related fields and per-row writes where a batch write exists, and warns on unbounded fetches and aggregation done in Python instead of the database. Where a pattern is deliberate, a # db-perf: allow comment with the reason suppresses the finding.

Two decisions stand out. Database performance findings can block legacy code too, the one deliberate exception to "never block legacy", because a query in a loop costs the same wherever it lives. And the workflow file did not change at all: adding a specialist meant one rule document, one agent definition and one row in the routing table.

What we would tell another team

  • Start with the output format and the severity definitions. They survived every version; the trigger and the architecture did not.
  • Make every finding cite a rule. It turns arguments about taste into a conversation about whether the rule is right.
  • Never flag pre-existing code. Nothing destroys trust in an automated reviewer faster.
  • Verify before you post. A second pass that checks the quote, the line, the diff and the rule is cheap compared with a wrong "request changes".

This connects to six feedback loops for AI coding agents: the review stage only improves the next change if its findings are precise enough to act on, and it guards against the problem in why generated tests miss bugs.

Claude Code review: common questions

Can Claude Code PR review run automatically on every pull request? Yes. It runs in GitHub Actions through Anthropic's official action, triggered on pull request events or by an "@claude" comment.

Claude Code review skill, command or agent: which should I use? They fit together: a slash command as the local entry point, sub-agents so each reviewer holds one set of rules, and a skill for the reusable procedure. To package ours as a skill, put the steps in a SKILL.md with the rule documents and finding format beside it, and keep the specialists and verifier as sub-agents. How to create Claude Skills walks through the structure.

Does it replace human review? No. It catches rule violations reliably when the rules are written down and cited, which is most of what a human reviewer spends time on, but it is weaker at judging whether a change is the right change. A human approval is still required on every pull request. Review effort is also the cost headline numbers hide: a 90% agent PR merge rate can still describe a weak workflow.

Where this leaves us

The system we run today is not clever. It is a playbook, a routing table, a handful of reviewers that each know one set of rules, and a verifier that refuses anything it cannot prove. Every piece exists because a simpler version failed in a specific way, and if you are starting now you can skip most of those failures.

The playbook itself, with our company-specific rules removed, is here: a Claude Code review template you can copy.


Originally published on nulltensor.com.

Top comments (0)