Coleman Parkes surveyed 300 senior engineering leaders — most of them running mission-critical C/C++ systems — on behalf of Undo, a root-cause-analysis tooling vendor, and published the results on October 7, 2026. The headline pair of numbers: teams spend an average of 9.8 hours a week producing code with AI agents, and 16.9 hours a week debugging it. Debugging alone now accounts for 42% of the average working week. Writing code got faster. Understanding it didn't, and the gap between those two numbers is the actual story.
The number underneath the headline number
The one I keep coming back to is 35%: that's the share of AI-generated code the survey says reaches production before the team has fully understood it. Not before it's tested — before it's understood by the people responsible for it. That's self-reported by the engineering leaders accountable for these systems, not an outside estimate.
The downstream numbers make it concrete. 81% of respondents had at least one production incident or outage tied to this in the prior six months, with 14% hitting it multiple times a month. 93% had a root cause misdiagnosed at least once because of an AI hallucination. 91% saw a test escape, a serious defect, or poorly optimized code reach production anyway. 80% said coding agents simply struggle with hard problems once a codebase gets complex enough — which is why roughly a third of teams in the survey now restrict agents to comprehension and debugging work only in the parts of the codebase that are already simple.
Why faster generation didn't produce faster releases
This is the part that should reframe how teams read their own velocity metrics: 79% of the surveyed leaders said agents generate code significantly faster. The same 79% said the overall release cycle is no faster than before AI agents arrived. The time saved writing gets spent somewhere else — Undo's CEO Greg Law described it as engineers losing "days trying to unravel what went wrong and why" with "code that's almost, but not quite right." A line of code that's 95% correct isn't 95% as useful as a correct one. It's a debugging session with better production values.
This matches the shape of the cost an autonomous agent ran up catching a bug a human finds in five minutes: the generation step looks cheap in isolation, and the real cost shows up one step later, in whoever has to verify, debug, or clean up after it.
What this actually changes about how I'd check AI-assisted work
Track time-to-understand separately from time-to-write, for a real sample of PRs, for one sprint. Most teams can tell you how fast an agent produces a diff. Almost none can tell you how long it took a human to actually understand what that diff does well enough to debug it later. If your ratio looks anything like 9.8:16.9, your "AI made us faster" claim needs a release-cycle number sitting next to it, not a lines-of-code or time-to-first-draft number.
Treat an agent's root-cause explanation as a hypothesis, not a conclusion — the same way you'd treat a guess from someone who wasn't actually present when the system broke. 93% of leaders in this survey had already been burned by exactly that failure mode: a confident, wrong explanation that someone acted on before checking it against actual logs or traces.
# A cheap habit that costs nothing and catches this specific failure:
# before merging a fix proposed by an agent, ask it to cite the
# specific log line, trace, or test output its root cause rests on.
# "I don't have one" is a more useful answer than a plausible guess —
# and a plausible guess with no citation is the thing to distrust most.
If you can't currently tell which parts of a shipped feature were AI-generated versus hand-written, you can't test any of these findings against your own team — you can only guess whether they apply. Tagging code by origin, even roughly, is the only way to find out whether your release cycle has the same shape as this survey's, or a genuinely different one.
The rule
Generation speed and verification speed are two different numbers, and this survey is one of the more concrete attempts to date to put real hours against the gap between them. An agent that writes code five times faster than a human doesn't make the code five times easier to understand — understanding doesn't compress the way generation does, and 16.9 hours a week is what that mismatch costs when nobody's tracking it separately.
I'm Rohit, founder of Krova Cloud — disposable sandboxes priced per minute, built so an agent's debugging and verification loop against a real environment doesn't turn into an unbudgeted line item. More debugging stories at debugly.dev.
Top comments (0)