DEV Community

Cover image for Stop reviewing agent code line by line
Tiago Granelli Ribeiro
Tiago Granelli Ribeiro

Posted on

Stop reviewing agent code line by line

Writing code isn't the slow part anymore. An agent ships a feature in minutes and fixes a bug just as fast. What takes time now is us reading what it wrote.

Nobody reviews a 10,000-line PR. Anyone who says they do is lying to their team or to themselves, because what really happens is we scroll, look at the tests, open a file or two and approve. And the person who does read all of it is holding back a whole team that could ship far more.

DHH said it out loud

DHH talked about this when he opened Rails World in September. He said 37signals stopped writing code by hand:

Writing code by hand is no longer an economically viable skill for most programmers at most companies.

According to people who covered the talk, he hasn't typed a line since March and now prefers English to Ruby. He rebuilt a backend in Rust without knowing Rust and presented that as an advantage:

I don't know any Rust at all. I consider that a feature.

A lot of people got angry. I think he only said out loud what most of us already do quietly, which is ship code we haven't read.

Reviewing agent code line by line has become theater. We keep doing it because we always did and because it feels like control, but at this volume review catches very little and costs the time of the most expensive people on the team. When you accept that agents will write most of the code, you also accept that you won't review the way you used to, and what's left to solve is how to trust what ships without having read all of it.

What I do instead

I skip the line-by-line review and put the effort elsewhere.

A rule has to break a command. Whatever is written in a prompt or an AGENTS.md is a suggestion, and an agent ignores suggestions when it's convenient. If a screen must not read the database directly, that becomes a lint error, and the error message says what to do instead. The agent reads it, fixes the code and runs again before I open the diff.

Tests need to run against the real system, with a real database and a real browser. A test that mocks everything tests nothing; it only confirms the mock returns what you told it to return. I want tests checking how many queries an operation makes and whether a page works with the keyboard alone, which are things review almost never catches.

A backup that has never been restored doesn't count as a backup. There has to be one I have restored at least once, and a deploy I know how to roll back, because that's what lets me ship fast without being afraid.

And the architecture has to stop the agent from doing something stupid. Layers that can't import each other, an API contract with one source, types that won't let you pass a user's id where an order's id goes. What the compiler refuses, I don't have to go looking for.

With that in place I review what the product should do, and I leave how it was built to the checks.

Where a great engineer spends the day

Bugs still get through, like they did when humans reviewed everything. But fixing became fast and cheap. A bug that used to wait three sprints to get priority now takes ten minutes of conversation with an agent, so holding code in a review queue out of fear of bugs no longer adds up.

Someone asked me recently what separates a good software engineer from a great one. For me, in 2026, it's where each of them spends their time. A great engineer spends it building quality gates and guardrails so strict that no agent can put something bad in the code. The engineer who still measures their worth by how many PRs they reviewed is falling behind.

If an agent wrote bad code in your project, that's on you. Your gates were weak and let it through. Fixing only that spot is like giving the project fever medicine: the fever is the symptom, and the cause is still there. So besides fixing it, I turn it into a rule, a test or a tighter type, and the same mistake doesn't come back with this agent or the next one.

The people who already understood this are the ones doing well right now. They use their time to improve the foundation of the project, and they ship far more than they used to.

What it costs

An agent takes much longer to deliver a feature in a repository like this, because it writes the tests and has to pass every check. If you want something fast and cheap for a prototype, it isn't worth it. For software someone will maintain, I don't know another way to use agents seriously.

I have a side project where I apply these ideas, Slopproof. It's only an example, and you shouldn't follow it to the letter. Some of its dependencies are betas or release candidates on purpose, so it isn't something to copy whole into production. It's there to show the practices working together, and you can take whatever makes sense for your own project.

Top comments (0)