Why deflection rate misleads
Deflection counts conversations that didn’t reach a human, which includes every customer who gave up. It reliably overstates value, and it is the number vendors lead with for exactly that reason. A system can post an excellent deflection rate while quietly damaging retention.
1. Resolution, confirmed downstream
Count an interaction as resolved only if the customer didn’t come back about the same issue within a sensible window. This one change usually cuts a headline number substantially, and what remains is real.
Repeat contacts are also the cheapest place to find what the AI employee is getting wrong, because the second conversation is usually explicit about what the first one missed.
2. Human minutes returned
Measure the time your team no longer spends, not the volume the AI handled. These diverge when the AI takes on work nobody was doing anyway, or when reviewing its output costs nearly as much as doing the task.
If reviewing the output takes as long as doing the work, you’ve automated the wrong step.
3. Quality against your own baseline
Sample the AI’s handled conversations and score them the way you score your team’s. The bar is not perfection — it is your existing standard. Teams often discover the honest comparison is more favourable than they expected, and occasionally that it is much worse in one specific category worth pulling back.
4. Coverage nobody was providing
Some of the value is work that simply wasn’t happening: after-hours responses, follow-ups that were dropped, research nobody had time for. This shows up as new pipeline or faster response times rather than as saved cost, and it is frequently the largest line once measured.
Count it separately from savings. Mixing the two produces a number nobody outside the project believes.
Then compare against the real alternative
The comparison is not against a perfect system, it is against what you would otherwise do: hire, outsource, queue the work, or let it go undone. An AI employee that handles a bounded job at your existing quality standard for less than the alternative is worth its seat, and that is the whole test.