DEV Community

The AI Prism
The AI Prism

Posted on Originally published at theaiprism.com

The AI Productivity Paradox: Why the Gains Haven’t Shown Up Yet

Originally published on The AI Prism


We’re using AI everywhere and measuring it nowhere.

Walk into almost any knowledge-work office in 2026 and you will find assistants embedded in the email client, the code editor, the CRM, the support queue and the slide deck. Ask the same organisation what any of it did to output per hour, and the room goes quiet.

That silence is the paradox. Adoption is close to universal at the individual level, spending is enormous, and yet the macro statistics that are supposed to register a productivity boom look stubbornly ordinary. Something has to give — either the tools do less than the demos suggest, or our instruments are pointed in the wrong direction.

The honest answer, laid out well in Bjorn Roche’s essay on the AI productivity gap and in the long Hacker News discussion it kicked off, is that both are partly true — and that neither is a reason to panic. This is what technology adoption has always looked like from the inside.

The Paradox, Stated Plainly

The shape of the problem is old enough to have a name. In 1987 Robert Solow observed that you could see the computer age everywhere but in the productivity statistics — a line the NBER literature on the productivity paradox has been unpacking ever since.

The 2026 version is tighter and faster. US Census Bureau survey work found firm-level AI use climbing from low single digits in 2023 to roughly 9% of firms by late 2024, and the Business Trends and Outlook Survey has kept tracking the curve upward since.

Meanwhile US labour productivity growth has run in the low single digits — respectable, not transformational, and well inside the range you would expect from a normal cyclical recovery, per Bureau of Labor Statistics series.

So the question is not whether there is a gap. It is which of several unglamorous mechanisms is producing it.

Measurement Lag: We Count the Wrong Things

Productivity is output divided by hours. Both halves of that fraction are hard to observe in knowledge work, and AI makes them harder.

If a support agent resolves the same number of tickets but each reply is clearer, output measurement records nothing. If a developer ships the same number of features with fewer defects, the defects that never happened do not appear in any denominator. National accounts are built to count widgets and billable hours, not avoided rework.

There’s a subtler problem: much of what AI produces is free. Generated images, drafts and summaries that would previously have been purchased or skipped entirely show up as consumer surplus, and GDP-based statistics famously undercount consumer surplus, as the Brookings work on measurement and intangibles has argued.

Roche’s essay makes a version of this point from inside a company rather than inside an econometrics paper: most firms deploying assistants never established a baseline, so they have no way to know what changed.

Diffusion Takes Years, Not Quarters

The second explanation is the one economists find most boring and most convincing. General-purpose technologies pay out slowly because the technology is the cheap part and the reorganisation around it is the expensive part.

Paul David’s canonical study of electrification showed factories took roughly 40 years to reap the full productivity benefit of electric motors — not because motors were bad, but because the payoff required abandoning the central-shaft factory layout and rebuilding the floor plan entirely.

Erik Brynjolfsson and colleagues formalised this as the “J-curve” of general purpose technologies: measured productivity falls first, because firms are pouring resources into intangible complementary investment — training, process redesign, data plumbing — that the statistics treat as cost rather than capital formation.

If that model holds, we are currently in the dip. The trough is not evidence of failure; it is the price of the reorganisation.

Task Automation Is Not Workflow Automation

Here is the mechanism that most directly explains why individual enthusiasm doesn’t aggregate into organisational output.

Assistants are extremely good at bounded tasks: draft this, summarise that, refactor this function, name these variables. Almost every credible study measures exactly that kind of task. The Brynjolfsson, Li and Raymond study of a customer support deployment found roughly a 14% average increase in issues resolved per hour, concentrated among less experienced workers.

But a workflow is a chain of tasks with handoffs, approvals and queues between them. Speeding up one link in a chain moves the bottleneck; it does not necessarily move the throughput. Anyone who has drafted a document in ninety seconds and then waited nine days for legal review knows this in their bones.

Amdahl’s law is a useful mental model here. If AI touches 30% of the work and makes that portion twice as fast, the end-to-end gain is about 18% at best — and that’s before the coordination costs of the new step.

The gains show up in the statistics only when the queue between the links is redesigned, and queue redesign is a management problem, not a model problem.

The Quality Offset: Busier, Not Better

The fourth mechanism is the least comfortable one. Some of the output AI enables is work that nobody needed.

When drafting becomes nearly free, the marginal cost of producing a document collapses — so more documents get produced. More documents means more reading, more review, more meetings about the documents. The volume of artefacts rises while the volume of decisions stays flat.

Open-source maintainers have been unusually blunt about this. Several major projects reported a surge in low-quality, AI-assisted submissions that consumed more reviewer time than they saved, with the curl project’s experience with AI-generated security reports becoming the canonical example of generated volume imposing a review tax.

This is a genuine negative externality: the producer captures the speed gain and the reviewer absorbs the cost. In an aggregate measure the two cancel, and the statistics record nothing.

Who’s Actually Faster — and Who Only Feels Faster

The most instructive recent finding is one that cuts against the tools’ own users.

A 2025 randomised controlled trial by METR on experienced open-source developers working in their own large repositories found that participants were roughly 19% slower when using AI assistance — while believing they had been about 20% faster. The METR study writeup is careful about its limits, but that perception gap is the single most important number in this debate.

The pattern across studies is consistent rather than contradictory. Novices and people working outside their expertise gain the most. Experts in familiar, complex codebases gain the least and sometimes lose, because verifying a plausible-looking suggestion costs more than writing the line yourself.

Self-reported productivity is therefore close to useless as evidence. Everyone feels faster, because the friction of the blank page disappears and the friction of review is diffuse and unmemorable.

The policy implication is uncomfortable for vendors. The people most likely to renew a licence are the ones who gained the least, because the speed they felt was real to them even when the throughput was not. The people most likely to churn are the experts who quietly lost time and noticed. Adoption metrics and value metrics point in different directions, and most dashboards only show the first.

If your organisation’s AI ROI case rests on a survey asking employees how much time they saved, you do not have evidence. You have a mood.

What Would Actually Close the Gap

Strip out the extremes and a moderate picture emerges from the credible studies. The strongest gains cluster in tasks that are text-heavy, low-stakes and easily verified. The weakest cluster in tasks that are context-heavy, high-stakes and expensive to verify — which describes most of what senior people are paid for.

Adoption itself is lopsided. Anthropic’s usage analysis found delegation-style use growing relative to collaborative use, and the Anthropic Economic Index shows adoption concentrating heavily in software and technical writing rather than spreading evenly across the economy.

The widely cited finding that a large majority of enterprise generative AI pilots produced no measurable P&L impact should therefore be read as a claim about pilots, not about the technology. Pilots without process change rarely move a P&L, whatever the technology. For a broader read on how this pattern plays out in employment rather than output, see The AI Prism’s look at what’s actually happening to jobs.

The pattern repeats across every general-purpose technology before this one. Electricity, the computer, the internet, each spent years as a disappointment in the aggregate statistics before the reorganisation caught up. The mistake was never to expect gains; it was to expect them on the deployment timeline rather than the absorption timeline. AI is behaving exactly as its predecessors did, which is the most reassuring and most ignored fact in the whole debate.

None of this means the tools are useless. It means the gains are conditional, and the conditions are organisational rather than algorithmic. That is good news, because organisations can be changed faster than physics.

None of the fixes are technical, which is precisely why they are slow.

Measure before you deploy. A two-week baseline of cycle time, defect rate and rework volume on one workflow is worth more than a year of adoption dashboards. Seat counts and token spend measure input, not output.

Pick one end-to-end workflow rather than sprinkling assistants across ten. Gains are only visible when the whole chain — including the approval steps and the handoffs — is redesigned around the new capability.

Then account for the review tax explicitly. If generated volume increases, reviewer capacity has to increase or throughput standards have to tighten; otherwise the saving quietly migrates from producer to reviewer and disappears.

Finally, be patient with the macro data. Electrification took decades and computing took roughly twenty years to register clearly. Expecting a general-purpose technology deployed at scale from 2023 to show up in national statistics by 2026 was never a reasonable timeline.

The Management Question Nobody Asks

Underneath all the measurement debate sits a simpler, more awkward possibility: the gap persists because most organisations never treated AI adoption as a change-management project in the first place.

Rolling out assistants is treated as a software rollout — buy licences, send a training email, watch the dashboard. Reorganising work around a new capability is a different activity entirely, and it is the one the productivity literature says actually moves the number. The tool arrives; the workflow does not.

This reframes the paradox in a useful way. The lag is not only a statistical artefact or a diffusion delay. Some of it is simply the ordinary cost of deploying a general-purpose technology without the complementary investment the theory predicts. The good news is that this is the one mechanism on the list a manager can actually do something about on a Tuesday.

The Bottom Line

The productivity gap is not proof that AI doesn’t work, and it isn’t proof that the statistics are broken. It is what the middle of an adoption curve looks like when the tooling has outrun the organisational plumbing around it — task-level speed with workflow-level friction, real gains for novices, unproven gains for experts, and a measurement apparatus that was never designed to see any of it.

The interesting question isn’t when the numbers will finally arrive. It’s whether the organisations currently counting seats and licences will have built the baselines they need to recognise the gains when they do — or whether they’ll still be measuring adoption and calling it impact?

References

Bjorn Roche — The AI Productivity Gap

Hacker News discussion — The AI Productivity Gap

Brynjolfsson, Rock & Syverson — Artificial Intelligence and the Modern Productivity Paradox (NBER w24001)

Brynjolfsson, Li & Raymond — Generative AI at Work (NBER w31161)

METR — Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity

US Census Bureau — Business Trends and Outlook Survey

US Bureau of Labor Statistics — Productivity Data

Brookings — Productivity in the Age of Artificial Intelligence

Anthropic Economic Index

Brynjolfsson & Hitt — Computing Productivity (NBER w7833)

curl project — on AI-generated security reports

The AI Prism — What Is Actually Happening to Jobs

The post The AI Productivity Paradox: Why the Gains Haven’t Shown Up Yet appeared first on The AI Prism.


Cross-posted from theaiprism.com — Cutting Through the AI Noise 🧊

Top comments (0)