The product was real: paying customers, a proper login, a billing integration, a dashboard, an API, a mobile-friendly front end. Two founders had built it in four months with a coding agent doing most of the typing, and it was better than most four month builds I've seen. Then one founder left, the other needed to raise money and stop coding, and I was asked to take it over.
I said yes, and the first thing I did was count. Sixty-one thousand lines of TypeScript, not counting tests, in a repository 140 days old. For comparison, the biggest product I'd built myself, with a team of four over a year, was around 45,000. This had been written at roughly ten times that rate. The rate itself isn't the problem. What it says about how much of the code a person had read is.
So I asked the remaining founder. His honest guess was that he'd read maybe a fifth of it carefully. The rest he'd reviewed the way you review a demo: does the feature work, does the screen look right, ship it.
The first bug
The first bug report after I took over was that some invoices had the wrong tax rate, only some of them. On a codebase I knew, that kind of bug used to take an afternoon: find where tax is computed, find the branch that picks the rate, find the condition that's wrong.
This one took a week, and bad code wasn't really the reason. Tax was computed in four places. There was a computeTax function in a billing module, which I found first and which was correct. There was a second one inside the invoice PDF generator, written on a different day for a different feature, with its own rate table that was three months out of date. A third, inline in the checkout flow, called the first one and then adjusted the result for a discount case. And a fourth, in a Stripe webhook handler, recomputed the tax from the line items to check the incoming amount, with rules subtly different from the other three.
Each one made sense on its own. Each was written by an agent asked to build a feature, which looked at what was immediately around it, didn't find a tax function in scope, and wrote one. Nobody had read enough of the codebase to know there were already three. The bug was that customers who got the PDF saw one number and customers who got the email saw another, and which was "wrong" depended on which of the four you thought was canonical.
Every bug I found in the first two months looked like that. The logic was rarely wrong. It was duplicated, and the copies had drifted apart.
What vibe coding actually produces
I want to be careful here, because the easy version of this post is "AI code is bad", and that isn't what I found. Line by line the code was fine. It was better named than a lot of human code, consistently formatted, with reasonable error handling and tests that passed. Sample any 200 lines and you'd think the team was solid.
What the process produced was a codebase with no shape. Human teams, even bad ones, build up a shared picture of where things live, because they have to read each other's code to work on it. That picture is what stops the fourth tax function getting written: someone on the team says we've already got one of those, use it. When the agent writes and the humans check the output, nobody has to read anything, so the picture never forms, and every feature gets built as if it were the first.
In this repository the symptoms were:
Four tax implementations, three date formatting helpers, two permission checks with different rules, and a utils directory with 90 files, eleven of them named some variant of format.
Twelve database access patterns. Some Prisma, some raw SQL, some through a repository class that existed for four tables and not the other thirty.
Environment variables read in 60 places, with three different fallback conventions.
A test suite of 2,100 tests with a 91 percent pass rate, where the failing 9 percent had been failing for weeks and were being ignored, because the features they tested had been reworked and nobody deleted the tests.
You can't blame any of that on AI. It's what happens whenever code gets written faster than it gets read, and the agent just made the writing half possible at a scale humans couldn't reach before.
The week I spent not fixing anything
After the tax bug I stopped taking tickets for a week and did what the founders never had time for, which was read it. Not all 61,000 lines, but every module boundary, everything that touched money, auth or the database, and every file over 300 lines. I made a map of what lives where, which of the duplicates is canonical and which are dead.
Then I wrote the map down in the repository, as the kind of file the agent reads at the start of every session. Tax is computed in billing/tax.ts and nowhere else. Dates are formatted with lib/format/date.ts. Database access goes through the repository layer, and here's how to add a table to it. These modules are legacy and must not be extended.
That file changed the agent's behaviour more than any prompt tuning could have. The next feature it built used the canonical tax function, because the file said one existed and where. The founders had been starting the agent with a fresh, empty picture of the codebase every session, and it did what a new contractor does with an empty picture: built whatever it needed right in front of it.
Deleting
The second month was mostly deleting. About 14,000 lines went, nearly a quarter of the codebase, and no feature was lost. The duplicates were folded into the canonical version one at a time, each with a test asserting the behaviour of whichever version customers had mostly been seeing. The failing tests were deleted or fixed.1
The agent did most of the typing for this as well. The difference was that every task started with the map, each task was scoped to one duplicate, and I read every diff, because that was the point of the exercise. My job was reading. The agent's speed was still useful, it just couldn't be the only thing that mattered.
What I would have done in the four months
By the standards of what they were trying to do, get a product in front of customers before the money ran out, the founders didn't do anything wrong. They managed it. But if I'd been in the room, these are the changes I'd have pushed for, and they're small, and none of them slows the agent down much.
Keep the map from day one: a file that says where things live, updated whenever something new gets added. It costs a minute per feature, and it's the one thing that stops the duplication, because it gives the agent the picture a team would have had.
Read one thing per feature. At that pace you won't read the whole diff, so read the one function that touches money, or auth, or the database. Reading the tax function on the day the second one was written would have caught it that day.
Delete tests that fail for more than a day. Nobody trusts a test suite that's 91 percent green, and a suite nobody trusts isn't doing its job. Either the test is wrong and it goes, or the code is wrong and gets fixed. "We'll look at it later" is how you end up at 9 percent.
Search before writing. Put it in the map file: tell the agent to look for an existing implementation before writing a helper. Agents do this when told. By default they don't, because by default they're trying to finish the task in front of them.
Budget the reading. If the agent writes 500 lines an hour and a person can read 150 with attention, the team is falling behind by 350 lines an hour, and that compounds. Either slow down the writing, or accept that the gap is a debt with interest and plan to pay it, which is what I was hired to do.
A number to watch
If a team working this way puts one metric on the wall, it should be lines merged per week divided by lines a human read with attention.2 Once the ratio drifts past three or four to one, the map is going stale, duplicates are being born, and the debt is growing at a rate whoever inherits the code will end up paying.
Looking back, the founders' ratio was somewhere around ten to one. Mine on the same codebase is now about two to one, and features ship at the same pace. The agent is spending its speed on the right things, because someone has read enough to know what they are.
Where it ended up
Six months on, the codebase is 52,000 lines with a map, one implementation of everything that matters, and a test suite that's green or the build fails. The agent still writes most of the code. The founder reads the diffs for anything under billing and auth, and I read the rest. Features ship about as fast as in the first four months, which surprised me. I think it's because the agent spends less time working around the mess it used to make.
What I keep coming back to is that none of this is new. A team that hired ten contractors in 2015 and never read their code would have ended up in the same place. What's new is that one person with an agent can now turn out what ten contractors did, so a reading deficit that used to take a big team to build up can be built up by a founder in a spare bedroom in four months, without them ever noticing. The tool didn't create the debt. It took away the friction that used to keep it small.
Originally published at zeybek.dev.
Top comments (0)