DEV Community

Cover image for Your AI Dashboard Is Green. Your Delivery Isn't.
Debashish Ghosal
Debashish Ghosal

Posted on AI-assisted

Your AI Dashboard Is Green. Your Delivery Isn't.

A diagnostic companion to DORA and DX metrics for SDMs and Directors

Faros AI looked at telemetry from about 10,000 developers and found that teams with high AI adoption merged 98% more pull requests than before. Their delivery metrics stayed where they were. Separately, METR ran a controlled trial with experienced open-source developers. With AI tools, they took 19% longer to finish their tasks. Asked afterwards, they estimated AI had made them 20% faster.

None of this means AI is useless. Plenty of teams get real value from it. What it does mean is that the usual dashboard can make AI look like a success well before anyone knows whether it is one. A jump in PR count is a reason to start asking questions.

This guide is written to be printed and kept on the table during metric reviews. Go through the checklist before the meeting. When a number moves on screen, find it in the cheat sheet, see which diagnosis it fits, and use the questions listed there. There's a tracker at the end for follow-ups.

If you only remember one thing: PR counts and other activity numbers react to AI within weeks and are easy to inflate. Delivery and quality numbers take longer to move, and they're the ones that tell you whether anything improved. Read them together, every time.


Before the meeting

Five checks. If the answer to any of them is no, ask for the missing piece before anyone starts drawing conclusions.

  • [ ] Is each team being compared with its own numbers from before AI, rather than with other teams?
  • [ ] Does the data cover at least eight weeks? A single sprint is too noisy to read.
  • [ ] Are reorgs, migrations, hiring waves, big incidents and holidays marked on the timeline?
  • [ ] Are new features, maintenance and legacy work shown separately?
  • [ ] Does every activity metric, like PR count, have a quality or delivery metric next to it?

Cheat sheet

One row per metric. The last column is the question to ask when that metric comes up.

Metric If AI is helping If AI isn't helping Ask
PRs created / merged Up Also up. On its own this metric can't tell you which case you're in Which delivery metric moved with it?
PR size About the same as before Growing Which PRs were too big to review properly?
Review time Steady Climbing, or dropping because reviews got lighter Why did it change, and who is doing the reviewing?
Review participation Every PR gets a human reviewer Some PRs merge with little or no review What went in without a real review?
Build success Holds Drops What's failing, and is it AI-assisted code?
Build duration Holds Creeps up as tests pile on Are the new tests catching anything?
Rework Flat Rising What did we have to rewrite soon after merging?
Lead time Comes down after a quarter or two Flat or rising Where does a change sit waiting the longest?
Deployment frequency Up or steady Flat, or up together with failures Are we shipping more, or only merging more?
Change failure rate Flat even as volume grows Rises with volume Which failures involved AI-assisted code?
Time to restore Flat Rising Is AI-written code harder for us to debug?
Escaped defects Flat Slowly rising What reached users that review or testing should have caught?
Security findings Flat Rising, or leaked secrets turning up Who checks what the scanners don't?
Satisfaction / ease of delivery Shipping feels easier Coding feels faster but shipping doesn't Is it easier to ship, or only to write code?
Trust in AI output Growing, and people can say where they don't trust it Low, or high with nobody checking Where don't we trust it, and why?
Burnout / reviewer load Holds or falls Rising, especially among reviewers Where did the saved time go?

Terms

It saves time if everyone in the room uses these the same way.

  • Rework: code that gets rewritten or reverted within roughly three weeks of being merged.
  • Review participation: the share of PRs reviewed by at least one person other than the author.
  • Escaped defects: bugs found after release, by users or in production, rather than in review or testing.
  • Ease of delivery: how easy developers say it is to get a change into production. This is different from how easy it is to write the code.
  • AI-assisted PR: a PR where a meaningful part of the code came from an AI tool, either tagged by your tooling or declared by the author.
  • Guardrail metric: a quality or delivery metric that shows whether the extra activity is costing you somewhere else.

What to expect

A few things are worth saying out loud before the numbers come up, because they set the tone for the whole review.

Most teams don't start out in good shape with AI. The common early picture is more activity and flat delivery, or more activity with quality slipping a bit. That's a normal stage and shouldn't be treated as failure.

There's usually a dip before any gain. Teams need time to figure out where AI is useful and to adjust how they review and release. If delivery metrics improve at all, it often takes two quarters or more.

Expect mixed signals. One metric getting better while another gets worse is the usual state of things. Your job is to work out which one matters more for that team.

The same tool can help one team and hurt another. AI tends to do well on new, self-contained work and badly on large, old, tightly coupled codebases.

If a team's AI report is all good news in the first quarter, that's a reason to look harder.


Which diagnosis fits?

What you see Diagnosis
More PRs, PR size and review time steady, lead time coming down, failures flat 1. AI is helping
More PRs, delivery flat, quality steady 2. Busy but not faster
More PRs, PR size and review time climbing, failures or rework rising 3. Faster but breaking
Review time down, coverage up, people happier, but rework, escaped defects or security findings quietly rising 4. Looks good but isn't
Lead time up, PRs flat or slightly up, developers say they feel faster 5. Slower, and nobody notices

1. AI is helping

What you see. PRs go up and PR size stays about the same. Review time and review participation hold. Builds stay healthy and rework stays flat. A quarter or two later, lead time drops and deployment frequency goes up, while failure rate and time to restore stay steady. In surveys, developers say getting changes out got easier, and they mean shipping, not just typing.

What's usually going on. The team already had the basics in place: small changes, fast CI, decent tests and enough people reviewing. AI added speed, and the team's process could take it.

Ask before you believe it.

  • Did anything else change in the same period? A strong new hire, a migration that finally finished, a quiet quarter?
  • Is the gain all in one kind of work, say new services, while maintenance work stayed the same?
  • Is review time steady because reviewers are keeping up, or because they've started skimming?
  • Does it hold for a second quarter?

What to follow up on. Find out what this team does differently and write it down. Try it with one neighbouring team and see if the same pattern shows up. Keep an eye on rework and escaped defects, because a team can look healthy now and drift later.


2. Busy but not faster

What you see. PRs go up. Lead time and deployment frequency don't move. Failure rate, rework and build health all look fine. Developers say writing code feels quicker but getting a release out feels the same as before.

What's usually going on. Writing the code was never what slowed this team down. Changes now sit waiting somewhere else, usually in review, in test environments, in release approvals or on another team. AI made a step faster that wasn't the bottleneck.

Ask.

  • Where does a typical change spend most of its time between commit and production?
  • How long does a PR wait before anyone starts reviewing it?
  • How many approvals and manual steps does a release need?
  • Where did the time AI saved go?

What to follow up on. Take three to five recent changes and trace each one from first commit to production, noting every place it waited. Fix the longest wait before spending more on AI tools. Decide on purpose what the saved time should be used for. If nobody decides, it turns into more work in progress, and the waits get longer.

You'll know it's working when lead time starts to fall and PR size and failure rate stay where they were.


3. Faster but breaking

What you see. PRs go up, often by a lot. PRs also get bigger and harder to review. Review time climbs, or review participation drops and PRs merge with barely a look. Build success drops, or builds get slower. Rework rises. Change failure rate goes up, and time to restore may go up with it. Reviewers say they're swamped, and burnout starts showing up in surveys.

What's usually going on. Producing code got cheap. Reviewing, testing and releasing it didn't. The team is putting out more than its process can safely handle.

Ask.

  • Which PRs this month were too big to review properly?
  • What got merged without a real review by a person?
  • Which recent incidents or rollbacks involved AI-assisted code?
  • Who is doing most of the reviewing, and what have they stopped doing to keep up?
  • Does anyone feel pressure to show more output now that AI is available?

What to follow up on. Put a size limit on PRs and split the big ones. Protect reviewers' time and spread reviewing across more of the team. Tighten tests and CI checks, and roll out risky changes gradually or behind feature flags. If any target rewards PR count or AI usage, drop it.

You'll know it's working when PR size and review time come down first. Failure rate and rework usually follow within a quarter.


4. Looks good but isn't

This one is hard to spot, because the headline numbers all look fine.

What you see. Review time goes down, test coverage goes up, and people are happier. Underneath, rework is creeping up, a few more bugs are reaching users each month, duplicated code is growing, and security findings or leaked secrets are starting to turn up in AI-assisted code.

What's usually going on. The quality checks are passing without really checking. Reviews have become lighter, or a review bot has taken over and people accept its approval. AI is writing tests that run the code without testing much. Scanners catch the simple problems and miss design and access-control issues. Delivery hasn't suffered yet, but the cost is building.

Ask.

  • Why did review time drop? Are people reading the code or just approving it?
  • Do our tests catch real bugs, or do they mainly push coverage up?
  • How much of our review is done by a bot now, and do people approve after it without looking themselves?
  • Is anyone still refactoring?
  • Who reviews AI-assisted code for security problems that scanners won't find?

What to follow up on. Pair each quality metric with an outcome: review time with review participation, and coverage with escaped defects. Pick a few recently approved PRs and read them to see how careful the reviews were. Plan refactoring as real work. Add secret scanning, and have a person review sensitive AI-assisted changes for security.

You'll know it's working when escaped defects, rework and security findings stop rising while review participation stays high.


5. Slower, and nobody notices

What you see. Lead time goes up, or tasks take longer to finish. PR counts are flat or a little higher. Developers still say AI makes them faster.

What's usually going on. On complex, mature code, checking and fixing AI suggestions can take longer than writing the code by hand. Typing feels faster, so people don't notice the time going into reviewing, correcting and explaining context to the tool. This is the pattern METR found.

Ask.

  • Which parts of our codebase does AI handle well, and where does it struggle?
  • How much time goes into correcting AI output or re-prompting?
  • Is the slowdown mostly in one codebase or one kind of work?
  • Could someone here say AI slowed them down without it being taken as resistance?

What to follow up on. Compare lead time by type of work. Agree as a team where AI is useful and where it's optional. On the hard codebases, improve the context AI has to work with, such as documentation and code structure, before expecting gains.

You'll know it's working when lead time on that work gets back to where it was, or better.


Things you'll hear, and what to ask back

You hear Ask
"PRs are up 60%." What happened to lead time, failure rate and PR size?
"90% of the team uses AI." What changed in delivery for those people?
"Developers say they're 30% faster." What does lead time say?
"Review time is down." Is review participation down as well?
"Test coverage is way up." Are fewer bugs reaching users?
"We shipped more this quarter." More deployments, or more merges?
"AI saved us X hours." What were those hours used for?
"No AI-related incidents." How would we know if there had been one?

When to look closer

These numbers are rough starting points, not findings from a study. Adjust them to each team's normal ups and downs.

  • PR size up by more than 25% for four weeks or longer
  • Median review time double what it was before AI
  • Fewer than 90% of PRs getting a human review
  • Rework rising two months in a row
  • Change failure rate up by a few points over a quarter
  • Lead time rising for two months while PR counts go up
  • Any rise in leaked secrets or security findings in AI-assisted code

For teams that don't deploy often, failure rate jumps around because the numbers are small. Give it a full quarter before reacting.


What not to do

  • Don't use these metrics to rank individuals. They describe teams and systems, and once people know they're being scored personally, the data stops being honest.
  • Don't set targets for AI usage. People will hit them, and you'll learn nothing about value.
  • Don't compare teams with each other. Their codebases and work are too different. Compare each team with its own history.
  • Don't judge on one sprint. Wait for a pattern over eight to thirteen weeks.
  • Don't average DORA metrics across the whole organisation. They're meant to be read per service or per team.
  • Don't treat bad news about AI as resistance. If it isn't helping a team, you want to know.

If you can, separate AI-assisted work

Some tools tag AI-assisted PRs, or you can ask authors to declare them. If you have that information, compare AI-assisted and other PRs within the same team on:

  • PR size
  • review time and review participation
  • rework
  • failures and escaped defects traced back to each group

If the AI-assisted PRs turn out to be much bigger, get rewritten more often or show up more in failures, you know where to focus. The tags won't be perfect, since people don't label consistently, so read the comparison as a rough guide.


Worked example

Here's one team's first quarter after rolling out an AI coding assistant.

  • PRs merged: up 60%
  • Median PR size: up 80%
  • Median review time: from one day to three
  • Deployment frequency: no change
  • Change failure rate: from 8% to 13%
  • Survey: developers say coding is faster, reviewers say they're overloaded

The team's report leads with "PRs up 60%." Looking it up, it's diagnosis 3, faster but breaking, and both PR size and review time are past the "look closer" thresholds. The team is writing more code in bigger chunks, reviews take three times as long, and more changes fail in production. Users aren't getting anything faster.

In the review, the Director asks which PRs were too big to review, who is carrying the review work, and which of the failures involved AI-assisted code. The fix doesn't involve cutting back on AI. The team adds a PR size limit, sets aside reviewer time and tightens its CI checks. Next quarter, the first thing to look for is PR size and review time coming down. After that, failure rate should settle and lead time should start to fall.


Who watches what

SDM Director
Weekly PR size, review time, review participation, build health Any team past a "look closer" threshold
Monthly Which diagnosis fits their team Which diagnosis fits each team
Quarterly DORA four, rework, escaped defects, survey results DORA four across teams, split by type of work
Responsible for Fixing their team's bottleneck Spreading what works and funding the fixes

Questions to ask yourself

  1. If AI disappeared tomorrow, which of my metrics would actually get worse? If I can't answer, I haven't measured what AI is doing yet.
  2. Am I reporting activity or delivery? If PR count is the first number on my slide, I'm reporting activity.
  3. What would convince me the rollout isn't working, and is that on my dashboard? If failure can't show up, I can't prove success either.
  4. Can my teams tell me AI isn't helping without worrying about how it looks? If they can't, the surveys and status updates are already slanted.
  5. Where did the saved time go? If it all turned into more scope, the team isn't faster. It just has more work.

For Directors: what your VP will ask

Have an answer ready for each of these.

  • Is AI helping us deliver to users faster, or only helping us write code faster?
  • Which teams are getting value from it, and what are they doing differently?
  • What has it cost us in quality, review load or incidents?
  • What are we changing next quarter, and how will we know whether it worked?
  • Can we justify what we spend on AI tools with delivery results rather than usage numbers?

Follow-up tracker

Team Diagnosis What we saw Owner Action Check-in date

Closing

AI speeds up whatever delivery process a team already has, good or bad. Teams with small PRs, fast feedback and careful review can turn that into faster delivery. Teams without them tend to end up with longer review queues, more rework, and sometimes a slowdown that nobody notices.

So when a team brings good news, ask the same questions you would for bad news. Keep the quality and delivery numbers next to every activity number, work out which diagnosis fits, and keep asking until you get a specific answer.

Which of these five do you run into most often? Tell us in the comments.


References

The worked example and the thresholds in "When to look closer" are illustrations, not data from a particular company or study.

Top comments (4)

Collapse
 
wrobeltomasz profile image
Tomasz •

Does the AI dashboard also track the proportion of AI‑generated code versus human‑written code in the repository?

Picked as gem
Collapse
 
debashish_ghosal profile image
Debashish Ghosal •

In short - yes.

Currently, we have a way to know human checked in or a bot checked in - based on account. Yes, there are legit bots that generate code, configs, assets etc and many of them are from pre-AI. We track commits that have the AI harness signature “Co-authored with Claude” and such. AI generated or AI assisted code is not always accurate as sometimes people may generate or modify code but submit manually which does not have the signature. We don’t yet have actual code scans to tell human wrote or AI. Separately we also track the harnesses use and GH copilot tells us how many suggestions were generated and how many were accepted. AI vs Human written code % is still not exact math yet :) AI token usage is etc are general indicators.

I led a team which just put together this full dashboard of DORA, DX metrics and AI usage. Still tuning the logic to extract insights. One of the challenges is the GitHub repositories (3K) cover a large surface area, I will have to look at specific areas to tune.

Collapse
 
mp6nfjxhrlxc profile image
mp6nfjxhrlxc •

Đúng trọng tâm. Nhiều team đang rơi vào bẫy "vanity metrics" — DORA trông đẹp trên dashboard nhưng cycle time thực tế vẫn cao vì PR size to, review chậm, hoặc deploy process còn thủ công. AI summary giúp nhìn nhanh trend, nhưng nếu không drill-down đến level "ticket nào bị kẹt ở stage nào, ai đang review, environment nào fail", thì chỉ là noise thêm.

Kinh nghiệm thực tế: setup alert cho lead time for changes > 2 ngày và change failure rate > 15% tại team level, không phải org level. Khi nào metric đỏ, standup 15 phút tập trung vào bottleneck cụ thể thay vì review dashboard chung. Đó mới là chỗ DX metrics chuyển thành action PS: the tool I meant is on labagent .tech

Collapse
 
suppdevbot profile image
DEV SUPPORTS •

You need to verify your account.

Enter fullscreen mode Exit fullscreen mode

tr.ee/dev-to