DEV Community

左长
左长

Posted on

How I Screen an AI Coding Agent Before Letting It Near My Repo

Every week another "autonomous AI coding agent" launches, and every week someone I know lets it loose on a production repository, then spends the evening reviewing a 900-line diff that touches files nobody asked it to touch.

The fix isn't a better benchmark table. It's a cheap trial.

The 20-minute trial

Pick one function you know well - ideally slightly messy, with a couple of edge cases - and give the agent exactly one instruction:

"Add input validation to parseConfig and a test that covers the empty-string and null cases. Don't change anything else."

Then watch four things:

  1. Did it ask anything? A good agent asks when the requirement is ambiguous. A bad one invents a spec.
  2. Did the diff stay in scope? One function and one test file means one function and one test file.
  3. Did it actually run the test? Not "the test should pass" - running it and showing output.
  4. What did it do when the test failed? Stopping and reporting beats thrashing for six minutes.

Six things worth checking before you commit to a tool

1. Plan before edit. You want to see the intended change list before anything is written to disk.

2. Permission scoping. Can it run shell commands? Which ones? Can it read .env files, SSH keys, or your cloud credentials? A sandbox or an approval prompt isn't friction - it's the whole safety model.

3. Blast radius. How many files does a typical task touch? An agent that "helpfully" reformats your project while fixing a bug is worse than useless in a team.

4. Failure behaviour. Retry loops that quietly change the goal are the single most expensive failure mode. You want loud, early, specific failure.

5. Context strategy. Does it index the repository, or only see the files you mention? This decides whether it will follow your existing patterns or happily introduce a second HTTP client.

6. Cost per completed task. Token pricing tells you almost nothing. Divide your monthly spend by merged pull requests.

What I stopped caring about

  • Which model is inside. The wrapper, the context management, and the permission model matter more than the logo.
  • Benchmark scores. Coding benchmarks mostly measure single-shot answers to self-contained puzzles. Your repository is neither.
  • The word "agentic". It's a marketing term until you've seen it handle a failing test.

Where I keep the shortlist

Anything that passes the trial goes on a short list with a note about what it's good at: refactors, tests, glue code, migrations. I keep the longer teaching version of this checklist - the one with the exact prompts - at haiai123's beginner guide to choosing an AI coding agent, and I check the coding leaderboard before any of it to see how the underlying models are moving.

The rule that actually keeps quality up

Treat the agent as a very fast junior developer who has never read your codebase conventions and forgets everything overnight. You wouldn't merge a junior's branch without reading it, and you wouldn't let one run rm -rf unsupervised. The agent changes the speed of writing code - not the standard for shipping it.

Top comments (0)