Originally published at ictdesk.net
Support teams running AI assistants are reporting deflection rates above 45 percent. Full self-service resolution, measured properly, sits closer to 14 percent. That gap is not a measurement quirk. It is roughly thirty one out of every hundred customers who stopped talking to your bot for a reason nobody recorded, and it is quietly the most flattering number in most support dashboards.
The two metrics are measuring different events
Deflection counts conversations that ended without a human agent. Resolution counts conversations where the customer got what they came for. Those overlap heavily, which is why the distinction survives so long unexamined, but they are not the same event and they fail in opposite directions.
A deflection counter fires on absence. The conversation ended, no agent was involved, increment the total. It has no view into whether the customer closed the window satisfied, closed it annoyed, or closed it and immediately opened your competitor's pricing page.
This is not a criticism of the people who built the metric. Deflection was a sensible proxy back when the alternative was a knowledge base article and the only question was whether someone read it. Applied to a conversational assistant that can hold a customer in a loop for eight minutes before failing, it stops being a proxy for anything useful.
Where the thirty one points go
The missing customers are not evenly distributed across mysterious outcomes. They cluster into four fairly predictable groups, and each one is visible if you go looking.
Some abandon partway through. They ask a question, get a partial answer, ask again in different words, and leave. The conversation shows engagement, multiple turns, and no escalation, which is the profile of a very good interaction and also the profile of a very bad one.
Some come back later. They give up on Tuesday afternoon and try again Wednesday morning, possibly reaching an agent that time. Two conversations, one problem, and the first one still counts as deflected.
Some switch channels. They leave the chat and open a ticket, or find a phone number, or reply to an old email thread. Your deflection metric records a success and your ticket volume records a new ticket, and no system connects the two.
And some simply stop. No follow up, no ticket, no complaint. These are the ones that never appear in any support metric at all, which is exactly why they are the expensive ones.
How your numbers compare
Published benchmarks for 2026 put median tier one deflection around 41 percent, with the top quartile near 59 percent. Those figures get quoted a lot in vendor material, and they are real. They are also measuring the avoidance number rather than the outcome number.
Roughly two thirds of AI support deployments come in below their deflection target in the first six months. The instinctive response is to widen what the assistant will attempt, because a bot that tries more questions deflects more of them. That reliably moves deflection up and just as reliably moves resolution down, since the added questions are the ones it was previously sensible enough to hand off.
I would rather see a team report 30 percent deflection with 26 percent resolution than 55 percent with 14 percent. The first team knows what their assistant can do. The second team has a better slide.
Four numbers that catch what deflection misses
None of these require new tooling. They require joining data you already have, which is a smaller project than it sounds and a less popular one than it should be.
Repeat contact within seven days. Same customer, same topic, second conversation. This is the closest thing to an honest resolution measure that most help desks can produce without changing anything. If someone comes back inside a week, the first interaction did not land, whatever it was counted as.
Channel switching after an AI conversation. Track whether a customer who ended a chat opened a ticket or called within the following hour. Today those two events live in separate tables and both look like successes. Joined on customer identity and timestamp, they become one failure.
Abandon point. Not just how many conversations ended without resolution, but where in the exchange they stopped. Conversations that die on turn two are a coverage problem, meaning the assistant never had the answer. Conversations that die on turn six are a comprehension problem, meaning it kept trying and could not get there. Those need different fixes and the aggregate number hides which one you have.
CSAT split by closure type. Score AI-closed conversations separately from agent-closed ones. Blended satisfaction is the single most misleading number in support reporting, because a strong agent team will float an underperforming assistant for a long time before anyone notices the average slipping.
The channel switching metric is usually the one that changes minds. Teams expect a small number and find that a meaningful share of their deflected conversations are followed by a ticket from the same person about the same thing within the hour.
The cost argument, handled honestly
Per interaction costs for AI-handled queries land somewhere near $0.62 against roughly $7.40 for a human-handled one. That ratio is real and it is why every support organisation is having this conversation.
But the comparison is only clean when the AI interaction actually ends the matter. An AI conversation that fails and produces a ticket costs $0.62 plus $7.40, plus a customer who is now on their second attempt and less patient than they were. Price the failures at their real cost and the arithmetic gets less dramatic, though it stays firmly in favour of automating the questions that automate well.
That is the more defensible version of the business case anyway. Automating high volume, low variance questions produces genuine savings that survive scrutiny. Automating everything produces an impressive deflection number and a slow leak in customer patience that shows up two quarters later in retention, where nobody connects it back.
Why so many pilots stall before production
Around 64 percent of enterprise CX teams have piloted agentic AI in support. Around 27 percent have anything in full production. That gap has the same root as the deflection gap.
Pilots are measured on deflection because it is the number available on day one. Production decisions get made on customer outcomes, because by then somebody senior has read the CSAT breakdown. A pilot that optimised hard for the first metric arrives at the production gate with results that do not support the case it was built on.
Teams that instrument resolution from the start have a less exciting first month and a much easier conversation at month six. They also know which categories to expand into, because they can see where the assistant genuinely closes issues rather than where it merely finishes conversations.
The other common blocker is data. An assistant can only resolve what it can reach, and a lot of pilots are quietly limited by knowledge that was never structured for retrieval. Poor coverage shows up as low resolution while deflection stays healthy, because the assistant is perfectly capable of producing a confident non-answer.
Getting the handoff right matters more than the model
Every one of the four metrics above improves when escalation works properly, which is a slightly deflating conclusion for anyone hoping the answer was a better model.
A customer who reaches a human quickly, with their conversation history intact and without repeating themselves, does not become part of the thirty one point gap. They become a resolved ticket that happened to start in chat. The handoff between assistant and agent is where most of the recoverable value sits, and it is consistently underbuilt relative to the effort that goes into the assistant itself.
ICTDesk runs live AI agent chat alongside the ticketing system rather than in front of it, so an escalated conversation becomes a ticket with the full exchange attached and the agent picks up mid-conversation. That design is less about the AI being clever and more about making sure a failed automation costs one interaction instead of three. Teams building auto-triage and routing on top of that get the same benefit further up the funnel.
What to change this month
Start with measurement rather than the assistant. You cannot tune what you are not observing, and the observation work is a week rather than a quarter.
Add repeat contact within seven days to whatever report your support leads already read. One number, same place, no new dashboard.
Join AI conversation logs to ticket creation on customer identity and timestamp. Look at the one hour window first.
Split CSAT by closure type and stop reporting the blended figure entirely.
Sample fifty conversations that ended without escalation and read them. Nothing in a dashboard substitutes for this and it takes an afternoon.
Once you have a resolution baseline, set targets on it and let deflection be a diagnostic rather than a goal.
The reading exercise in step four is the one people resist and the one that changes behaviour. Every team that does it finds at least one category where the assistant is confidently wrong, and finds it faster than any metric would have surfaced it.
Frequently asked questions
What is a realistic self-service resolution rate?
Around 14 percent across mixed query types is a reasonable current benchmark, and teams with narrow, well documented product surfaces do considerably better. The useful comparison is against your own baseline over time rather than against an industry figure covering wildly different support workloads.
Should we stop reporting deflection entirely?
No, but demote it. Deflection is a good diagnostic for capacity planning and a poor one for customer outcomes. Keep it next to resolution so the pair can be read together, and the moment they start moving in opposite directions you have learned something.
How do we measure resolution without asking every customer?
Use repeat contact as the proxy. If the same person does not return about the same issue within seven days, treat it as resolved. It undercounts silent dissatisfaction, which is a known limitation, but it is far closer to the truth than deflection and it needs no survey.
Does a higher deflection rate always mean lower resolution?
Not always, but it usually does past a certain point. Deflection gains from better coverage of questions the assistant genuinely handles will lift both numbers. Gains from widening scope into questions it should be escalating will lift one and depress the other, which is why the pair needs watching together.
How long before a new AI assistant produces meaningful numbers?
Deflection stabilises within a few weeks. Resolution needs at least a full repeat contact window plus a few cycles, so realistically six to eight weeks before the figure means anything. Judging a deployment at week three will tell you about coverage and nothing about outcomes.
Is agentic AI different from a chatbot for these metrics?
The metrics apply the same way, but the failure mode shifts. A scripted bot fails early and visibly, so customers escalate fast. An agentic assistant fails late, after several plausible turns, which produces longer conversations and a larger silent abandon group. Better technology, harder to measure.
Top comments (0)