AI coding agents write fluent, syntactically clean code that compiles on the first attempt. That fluency is precisely why reviewing it is tricky: humans tend to equate smooth phrasing with correctness. In practice, agent-generated code fails in remarkably predictable patterns. It skips edge cases, hallucinates parameters on older SDKs, introduces subtle N+1 queries, and writes tautological tests that pass without actually verifying anything.
Reviewing agent output with vague impressions or casual skimming does not work. You need a fixed, repeatable checklist that targets the specific blind spots agents exhibit most often.
Here are eight checks across five critical axes, along with the concrete prompt to fix each one.
Axis 1: Correctness
1. Empty Collections and Boundary Handling
-
The Mistake: Agents default to the happy path. When iterating, slicing, or computing aggregate values, they rarely account for empty arrays,
nullpayloads, or single-element collections. -
Correction Prompt: "Refactor
calculateInvoiceTotalsto handle empty arrays, null line items, and negative quantities without throwing."
2. Plausible but Nonexistent API Signatures
- The Mistake: Agents frequently invent convenience methods or pass modern flags into older library versions that your project actually pins.
-
Correction Prompt: "Verify whether
client.fetchWithRetry()exists in our pinned SDK version; if not, rewrite it using standardfetch()and our internal retry utility."
Axis 2: Security
3. String-Interpolated Queries and Commands
- The Mistake: When combining dynamic filters or calling system utilities, agents often revert to string template literals instead of parameterized queries or argument arrays.
-
Correction Prompt: "Convert the string interpolation in
getUserRecordsinto parameterized SQL using indexed positional placeholders."
4. Permissive Defaults and Fallback Credentials
-
The Mistake: To ensure sample code works immediately, agents often insert placeholder fallback secrets (
process.env.KEY || 'default_secret') or default permissions to0777or wildcard CORS headers. -
Correction Prompt: "Remove all hardcoded fallback secrets in
authConfig; fail fast at startup if required environment variables are unset."
Axis 3: Readability & Architecture
5. Premature Abstraction Sprawl
- The Mistake: Given a straightforward requirement, an agent will often generate unnecessary interfaces, factory classes, and strategy wrappers for what should be a ten-line pure function.
-
Correction Prompt: "Inline
DataTransformerFactoryand replace the strategy classes with a single pure mapping function intransformers.ts."
Axis 4: Performance
6. Hidden N+1 Queries and Unbounded In-Memory Filtering
-
The Mistake: Agents routinely execute database calls or network requests inside loops, or pull an entire database table into memory only to filter it with
.filter(). -
Correction Prompt: "Move the query outside the
forloop insyncProfilesand batch it into a singleWHERE id IN (...)operation."
Axis 5: Test Quality
7. Assertions That Verify Nothing
- The Mistake: Agents optimize for getting tests to pass. They frequently write tests that only verify an HTTP 200 status code, without checking response bodies, updated database state, or emitted events.
-
Correction Prompt: "Update
testUpdateUserto assert the updated database row values and response payload fields, not just the HTTP status code."
8. Mocking the System Under Test
- The Mistake: To avoid complex setup, agents sometimes mock the internal business logic of the exact function they are supposed to be testing, creating an entirely circular pass.
-
Correction Prompt: "Remove the mocks on
calculateTaxRatewithin its own test suite and verify the actual calculation logic against fixed test fixtures."
The Output Protocol
When using an agent to review code—or when formatting your own review comments—enforce a strict output rule:
- Findings First: Order every finding strictly by severity (Critical, Major, Minor). Do not bury blockers under complimentary opening remarks.
-
Anchor to Source: Every comment must include an exact file path and line reference (
path/to/file.ts#L42-L48). If you cannot point to the line, do not file the comment. - No Invented Problems or Style Nitpicks: Prohibit speculative feedback and leave formatting to automated linters.
- Cap Diff Size: If an agent submits a diff exceeding 400 lines, reject it immediately and require the agent to split the work into smaller, isolated PRs before reviewing.
The complete checklist prompts and operational review skills are open-source and free to read on GitHub at https://github.com/alapha888/agent-skills-en . If you prefer a condensed, printable reference for your desk or tablet, a PDF and PNG bundle containing this code-review checklist alongside a conventional git commit cheat sheet is available on Gumroad at https://alapha888.gumroad.com/l/zjayxe for $19. It contains the same material formatted for quick visual reference while reviewing pull requests.
Top comments (0)