AI saves real time on individual tasks such as drafting, summarizing and boilerplate, but most of that saving never reaches anyone’s calendar: 2026 field data shows it leaking away through verification, downstream cleanup and heavier workloads. In Danish payroll records, chatbot adoption left earnings and recorded hours statistically unchanged, and workers’ own estimate of the saving averaged about 2.8% of their work hours. nberarxiv
TL;DR
Task-level savings are real where output is easy to check (drafts, summaries, boilerplate). They shrink or reverse where checking is expensive (spreadsheet analysis, ambiguous judgment calls).
The best-designed developer trial found a 19% slowdown in early 2025, while participants believed they were faster. The 2026 rerun was too compromised to give a current number.
Three leaks separate “time saved” from “time released”: verification, downstream cleanup and scope expansion.
About 89% of executives in a survey of nearly 6,000 firms report no labor-productivity effect from AI over the past three years. That is perception, not measurement. NBER
Stop tracking “hours saved.” Track time to accepted output and rework hours.
The study that broke because people liked the tool too much
In February 2026, METR, a nonprofit that evaluates AI systems, said it was redesigning its developer productivity experiment. The reason was awkward. More developers were declining to participate because they didn’t want to work without AI, and 30% to 50% of them said they held back tasks they didn’t want to do without it. METR
That is a strange kind of evidence. Nobody demands to keep a tool that slows them down, so the refusals look like a signal that AI helps. But METR’s earlier trial showed how unreliable that signal can be. Its 16 experienced developers forecast a 24% speedup, reported feeling 20% faster afterward, and were measured at a 19% slowdown across 246 tasks. METR’s later survey work puts the average overestimate at more than 40 percentage points. arxivmetr
This matters now because licence budgets and staffing plans are being set on “hours saved” figures. Those figures come from at least six measurement methods, and their answers run from a 19% slowdown on a developer’s own repository to an 80% speedup on tasks sampled from real Claude conversations. Nothing is wrong with the arithmetic. The methods measure different things, and most of what follows is about which thing each one measures. anthropic
Why the numbers disagree: each one measures a different rung
Picture a ladder of measurements: the speed of a task, the value of a task, one person’s output, one firm’s output, and finally payroll and hours. Every rung up drops information. A tool can make a task faster without making the task more valuable, and a more valuable task does not guarantee a firm produces more.
Here is how the main sources rank, from strongest to weakest evidence for the question “does AI free up working time?” Read More....
Top comments (0)