DEV Community

Cover image for Your AI Productivity Number Is Right. So Is Your Team’s.
Filipe
Filipe

Posted on Originally published at levelup.gitconnected.com

Your AI Productivity Number Is Right. So Is Your Team’s.

Ask your engineers how much AI has improved their productivity. You'll hear somewhere between fifty and a hundred percent. Transformative. Game-changing. Then ask yourself whether you could stand up in front of your leadership and say the same thing with confidence.

Most engineering managers can't - not yet. The gains feel real, but they're hard to pin down. And that gap between what your team experiences and what you can actually defend is more useful than it might seem.

That gap isn't a measurement failure. It's the most useful signal you have.

The measurement problem

The gains your engineers are reporting aren't exaggeration. They're just measuring the wrong thing.

Ask where AI has actually helped and you'll hear the same answer across most teams: automation tests and boilerplate code. The mechanical layer - the work that needs doing but rarely needs thinking. If a test suite that used to take an hour now takes ten minutes, that's an 83% reduction on that task. That's real.

But it's only half the picture.

Where the saved time actually goes

Laura Tacho at DX puts the industry average at 3.75 hours saved per developer per week. Take that on face value. The question isn't whether the time is being saved. It's where it goes.

The time reclaimed from the mechanical layer doesn't convert directly into new features.

AI-generated code creates its own overhead. PRs arrive faster, which means review pressure arrives faster too - and reviewing code you didn't write, at higher volume, is its own drain on your engineers' attention. Generated output varies in quality. Security and maintainability issues that an experienced engineer would have caught while writing don't always surface until someone else is reviewing.

Meanwhile, the work itself doesn't hold still. A ticket is a snapshot - a fixed description of a feature the Product Manager wrote before anything was built. Once she can see intermediate output, her understanding shifts, and so does the scope, and the time to build the feature increases.

That isn't dysfunction; it's how good product thinking works. The work that exits a sprint rarely resembles the work that entered it.

What to measure - and what to fix

Your delivery metrics give you something defensible for leadership - a number you can stand behind when someone asks what AI is actually delivering. Survey data gives you the signal from inside the work, and the frame to make sense of the gap between the two.

To close that gap, you need to invest in the engineering harnesses that make AI-generated code production-worthy without consuming the time AI just freed up. That means reducing the review burden from agent-generated PRs (better tooling, clearer standards, automated quality gates, and early detection of architecture drift). It means ensuring generated code is maintainable, secure, and consistent with your architecture before it reaches a human reviewer.

Your engineers reporting 50–100% aren't describing where you are. They're describing where you could get to, once the engineering discipline around AI catches up with the adoption of it.

And when you get there, you might find speed wasn't the real prize. Working code earlier means product thinking catches up, requirements shift, and you deliver the right thing. Your delivery metrics won't show that - but your business will feel it. That's the gap worth closing.


Filipe Albero Pomar is an engineering manager and speaker based in London. He writes about AI adoption, engineering leadership, and software delivery. Find him at alpomar.dev.

TechLeadConf 2026 · GenAI London 2026

Top comments (0)