A few weeks ago I ran into one of those bugs that looked simple enough that I kept thinking the next fix would probably solve it. A small Node.js and TypeScript service was pulling jobs from several APIs, normalizing them and storing them in PostgreSQL. The background worker finished without errors and logged that the import had completed successfully, but when I checked the database, some of the jobs simply weren't there.
Running the same import again would often make more of them appear. Not all of them, just more, which made the problem look like some kind of timing issue. So I asked the AI to take a look. 😎
The first suggestion was to add retries around the database writes. That seemed reasonable, but it didn't help. Then came a delay, changes to the transaction boundaries, a higher retry count and eventually some extra error handling around individual inserts. Four rounds later, the symptom was exactly the same, while the diff had spread through the worker, the database layer and the tests.
The actual bug turned out to be an async callback inside a forEach. None of the suggested fixes had touched it.
I don't think the AI was being dumb here. I think I kept handing it the same information and asking for a different answer.
Look at the symptom from the AI's side for a second. Rows go missing and a rerun brings more of them back and nothing in the description says where they went, so it smells like timing, right? Every fix it suggested followed that smell from the first round to the last.
First it added retries around the database writes. Then came a delay, different transaction boundaries and a higher retry count. By the last round, it was catching individual insert errors so one bad record wouldn't affect the rest of the batch.
Those guesses weren't unreasonable. If all you know is that some rows go missing and a rerun brings some of them back, a timing issue is an obvious place to look. I probably would have started there too. 🙄
The funny part is that there really was a timing problem, just not the one we were trying to fix. The writes weren't happening too early or taking too long. The worker simply wasn't waiting for them to finish.
And that's where the debugging nightmare really started, because by that point neither the AI nor I seemed to understand the problem any better than when we began. 😕 By the fourth attempt, the suggestions had started circling back to the same ideas, just with slightly different settings. The last one went in another direction, but it still didn't solve the problem. If an insert fails and the code catches the error and continues, the row is still missing; now the failure is just quieter.
Here's what I think was going on. Every round the AI got the same clues (some rows are missing and a rerun brings back more). Unless a model gets logs or a way to run things, it works from the description and the code it can see, and "still missing" adds nothing to either of those. So it guessed again from the same clues, except that now the conversation also held every guess that had already failed, and every new answer had to work on top of the code the previous answers had left behind in the worker and the database layer.
At some point I realized we weren't really debugging anymore. We were just generating new guesses from the same two facts: some rows were missing, and running the import again brought some of them back.
And every failed fix made the next round a little messier. Now the AI wasn't only looking at the original code. It was also looking at retries, delays and transaction changes that had been added during the previous attempts. The code was changing, but my understanding of the bug wasn't.
Could a better prompt have helped? Maybe, but I'm not sure what I could have put in it. Everything I knew was already in the conversation. I could rewrite "some rows are missing" five different ways, but it was still the same piece of information.
What I actually needed wasn't a better way to describe the symptom. I needed one new fact about what the program was doing.
The Fact Was In The Order Of The Logs
So I stopped asking for another fix, reverted the changes and went back to the worker as it was before all of this started. This time I followed a single import through the logs from beginning to end, including the database writes.
That showed something off right away, just from the order of the lines. The worker logged "completed" before all of the inserts had finished. A job can't be done while its own writes are still running, unless nobody's waiting for them.
Here's the shape of the code. On a quick read it looks fine:
jobs.forEach(async (job) => {
await saveJob(job);
});
console.log("Import completed");
forEach calls the callback once per job and ignores whatever the callback returns. An async callback hands back a promise as soon as it reaches its first await, so forEach ends up with one promise per job, drops every one of them on the floor and returns as if the work were done. MDN says it plainly: forEach() "expects a synchronous function" and "does not wait for promises." So the log line runs while the inserts are still in flight and the job gets marked complete with its writes still going in the background, and by the time some of those writes failed the worker had already reported the whole import as a success.
I think that's also why a rerun kept bringing more rows in. Every run got a few more of the inserts through.
The fix is to give the code something to wait on:
await Promise.all(
jobs.map((job) => saveJob(job))
);
console.log("Import completed");
Now the log line waits for every insert and if one of them fails Promise.all rejects and the code after it never gets to say "completed". It does start all the inserts at once, though, and MDN notes that a rejection "does not cancel the remaining operations", so after one failure the rest keep running. If you want the result of every insert, Promise.allSettled waits for all of them. And if parallel inserts aren't wanted at all, a plain loop does one at a time:
for (const job of jobs) {
await saveJob(job);
}
console.log("Import completed");
It's slower, but nothing's running behind your back. (There's also a lint rule for this exact mistake. typescript-eslint's no-misused-promises flags an async callback passed to forEach and it's switched on in the recommended-type-checked config, so a type-aware lint setup would have pointed straight at it.)
Anyway, once I saw the order in the logs, finding the actual bug took a few minutes.
After Two Failed Fixes, I Go Find A Fact
My rule now: if two AI fixes fail and I still haven't learned anything new about the bug, I stop asking for another fix. First I find one new fact: a repro, an error, a payload, a stack trace or just the real order of events like this time. Then the AI comes back in, with that fact. Not before.
What made me stop this time was the diff, honestly. It had spread into the worker and the database code and the retries and the tests while the symptom stayed exactly where it was, and on top of that one of the new suggestions was just a variation of something it had already tried. Those two together are a pretty reliable sign that the guessing's run out. I'd add one more: a fix that swallows errors has stopped looking for the cause.
I'm not saying AI is bad at debugging. I still use it for this all the time. But if I'm on the third fix and I still haven't learned anything new about the bug, I've probably stopped debugging and started guessing.
That's what happened here. A fifth fix would have been working with the same information as the first four. The logs were the first thing that actually changed what I knew.
How many failed fixes do you usually let an AI try before you stop and go back to investigating the bug yourself?
Top comments (2)
I don't realy measure it by a number of failed fixes. Instead, I stop the moment I realize the problem hasn't been decomposed to fit the tool's capability boundary.
Complexity isn't necessarily something the machine should swallow whole. Decompose the problem until each unit fits the capability boundary, then compose the results.
Did we actually give the tool a problem that fits the way the tool operates?
Complexity is also relative. If Hello World is complex to me, then it's complex in my problem space. I don't get to invalidate that because someone else considers it trivial.
That's why I think in particles. I wouldn't tell a child to cook soup. But I can ask them to bring the ingredients, then the next thing needed. Small, capability-aligned tasks can still contribute to a larger outcome.
And there's an important distinction: tools can still fail within their boundaries. But we should distinguish tool failure from misuse of the tool.
Care to read more? POC
I'd make the first step a test that reproduces the missing rows. Then each proposed fix has the same failure to work against, and the next session can pick up from that evidence.