← Blog/For Enterprise

What Is Ticket Deflection? Formula, Benchmarks, and What It Hides

Zendesk's deflection formula returns a ratio. Decagon's returns a percentage. Here is what the metric measures, and the four numbers that tell the truth.
Last updated: August 28, 2026
10 min read time. Summarize with:
What Is Ticket Deflection? Formula, Benchmarks, and What It Hides

Ticket deflection is the number AI support vendors lead with, and it is the number I trust least.

Not because it's fake. Because it isn't one number. Two of the highest-ranking definitions of "deflection rate" on the internet right now don't even return the same kind of value. One gives you a ratio. The other gives you a percentage. Both call it the deflection rate.

So when a vendor tells you their agent deflects 70%, you have learned almost nothing until you ask three follow-up questions. This piece covers what the metric measures and the two formulas actually in circulation. Then the published benchmarks by industry, and the four numbers I'd track instead if I owned the deployment.

What Ticket Deflection Actually Measures

Ticket deflection is the share of inbound support contacts that never become a ticket handled by a human agent.

An interaction usually gets counted as deflected when one of these happens:

  • The customer's question is answered and they don't escalate
  • The customer closes the chat or ends the call without asking for a human
  • The customer is routed away from the queue before a ticket is created

Read that last one again. A customer who abandons a chat because the bot gave them something useless has been "deflected" under most measurement conventions. Their problem is not solved. They are going to call back, email, leave a review, or quietly stop being a customer.

That's the whole issue with the metric. It counts an exit, not an outcome. Deflection can be optimized in ways that actively damage the business. A vendor who reports it without CSAT, resolution rate, or re-contact rate alongside it is showing you a third of the picture.

The Formula Nobody Agrees On

Here's the part that surprised me when I went looking. There is no single deflection formula, and the disagreement isn't cosmetic.

Zendesk publishes this one:

Ticket deflection rate = total users of your help center ÷ total users in tickets

Their worked example: if four customers resolve their issue in self-service for every one who files a ticket, your deflection rate is 4.

Decagon publishes this one:

Deflection rate = deflected contacts ÷ total contact attempts × 100

That returns 80% on the same underlying traffic.

Four and 80% describe the same support operation. They are not the same number, they are not on the same scale, and neither vendor is wrong. A third convention, common in help-center analytics, counts sessions rather than contacts: (self-service sessions − tickets created) ÷ self-service sessions × 100. A fourth only counts a contact as deflected if the customer didn't come back about the same issue inside a 24 to 72 hour window. That one is the strictest, and the most honest.

So the first question to ask any vendor quoting a deflection rate is not "what's the number." It's "what's the denominator." The second is "what window." The third is "does an abandoned session count."

Most demos do not survive the third question.

Deflection vs. Containment vs. Resolution

These three get used interchangeably. They measure progressively harder things.

  • Deflection rate: contacts handled without a human, regardless of outcome. Includes abandonment, includes customers who gave up, includes sessions the bot simply closed.
  • Containment rate: contacts where the customer reached the end of the AI interaction without escalating. Higher bar, because it excludes abandonment. Still doesn't confirm the customer got what they needed. Decagon frames containment as the channel-level version of the same question, which is worth knowing when you're comparing two vendors' dashboards.
  • Resolution rate: the customer's issue was actually solved. Measured through post-interaction CSAT, re-contact rate, or an explicit confirmation step in the conversation.

Track all three. The gap between them is the most useful diagnostic you have. High deflection with low resolution means your agent is good at ending conversations and bad at finishing jobs. That specific gap is what customer self-service programs have been quietly producing for a decade, and it's why the metrics you pick matter more than the platform you buy.

Evaluations Docs
Voiceflow lets you score every transcript against criteria you define, including deflection, resolution and CSAT on the same conversation.

What Deflection Rates Are Realistic

Published bands, not vibes. Decagon's glossary gives per-industry deflection ranges that match what I see in practice:

IndustryPublished deflection band
E-commerce55% to 75%
SaaS40% to 60%
Financial services and healthcare25% to 45%

The spread tracks interaction complexity and regulatory load, which is the right thing for it to track. A returns question and a disputed transaction are not the same job.

Now set those against the two numbers everyone quotes. Gartner predicts that agentic AI will autonomously resolve 80% of common customer service issues by 2029. The present-day figure from the same firm: roughly 14% of issues fully resolve in self-service today.

Both are real. One is a forecast for 2029 and one is where the average team is standing, and the 2029 date falls off in almost every retelling. We wrote a whole piece on what closes that gap, and the short version is that it's content coverage and integration depth, not model choice.

Which means: if your vendor is quoting 80% to 90% deflection across all interaction types in month two, one of three things is true. The interaction mix is unusually simple. The deflection definition includes abandonment. Or the number isn't being measured the way you think it is. Ask for re-contact rate and CSAT from the same interaction set and watch what happens.

What Klarna's Numbers Actually Showed

The best available case study is public, and it went both ways.

In February 2024, Klarna announced an OpenAI-built assistant handling two-thirds of customer service chats, described as the work of 700 full-time agents, with resolution time down from 11 minutes to under two. I went and tested that agent once the fanfare died down and published what I found: responses that ran three times longer than they needed to, and several that were flatly wrong.

On 8 May 2025, CEO Sebastian Siemiatkowski told Bloomberg the company had gone too far. His words: "We focused too much on cost. The result was lower quality." Klarna resumed hiring human agents and committed to always offering one.

Nothing about the two-thirds figure was false. The agent really did handle that volume. The deflection number was excellent and the outcome underneath it wasn't. The company with the most to gain from the headline is the one that published the correction. If you need a single argument for why deflection alone shouldn't be a target, that's it.

See how leading teams design, test, and deploy AI agents at scale.

The Four Metrics That Actually Tell You If It's Working

If deflection is a misleading headline number, here's what goes on the dashboard instead.

1. Successful Containment Rate

Contacts fully handled by the AI without escalation and with a positive or neutral outcome signal. Pair containment with post-interaction CSAT or a resolution confirmation step at the end of the conversation. This is your primary performance metric, and it's the one that should be in the QBR deck.

2. Re-Contact Rate

The share of customers who come back within 24 to 48 hours of an AI-handled interaction. A high re-contact rate means the AI closed conversations without resolving them. This is the single clearest signal that a deflection number is inflated, and it's cheap to instrument. Track it from week one.

3. Escalation Quality

When the agent does hand off, how good is the handoff? Does the human have the context? How fast does the escalated interaction close? A bad escalation doesn't just fail to contain. It makes the human-handled contact more expensive than it would have been without the bot in front of it.

4. Cost Per Resolution

Fully loaded cost of resolving a contact, blended across AI and human handling. This is the number that connects agent performance to the business case. As successful containment rises and your humans handle a more concentrated set of genuinely hard interactions, cost per resolution should fall. If it isn't falling, something in the deployment is off, and the deflection rate will not tell you what.

How to Measure Those Four Without a QA Sample

Here's the operational problem with everything above: traditional QA is a spot check. Someone reads a handful of transcripts, maybe 5% of conversations, and the other 95% is hope.

That's fine for catching an agent behaving badly. It is useless for computing a resolution rate, because a rate calculated on a 5% sample is a rate with a confidence interval nobody puts in the deck.

Evaluations solve this by scoring every transcript against criteria you define, automatically. Resolution, deflection, CSAT, escalation appropriateness, and citation compliance can all be criteria on the same conversation. That is what lets you compute the gap between deflection and resolution instead of arguing about it. In Voiceflow, evaluations return binary, rating, or freeform results. "Did the agent fully resolve the issue" and "rate customer sentiment 1 to 5" live in the same run. Observability then aggregates the failures so you can see the pattern instead of five anecdotes, and the reporting layer turns that into something a VP can act on.

Teams running help desk automation or broader customer service automation programs tend to arrive here eventually, and so do teams still standing up their first support agent. It's faster to arrive on purpose. Turo, StubHub, Sanlam and Trilogy all run this way.

Why Containment Climbs Over Time

The teams with the best containment at 12 months are rarely the ones with the best containment at 30 days. Early numbers are limited by knowledge gaps, integration scope, and calibration, and all three are fixable.

The drivers are predictable:

  • Knowledge coverage. Most agents launch with partial coverage of what customers actually ask. Reviewing failed transcripts and filling the gaps moves containment more than any model change will.
  • Integration depth. An agent that can take an action keeps customers that an agent which can only answer questions loses. Every new system you connect, order management, billing, returns, converts a category of contacts from acknowledged to resolved.
  • Escalation tuning. Early deployments escalate too eagerly or, if someone is being measured on deflection, not eagerly enough. Tuning the triggers improves both containment and the experience of the people who genuinely need a human.
  • Failure pattern analysis. At volume, aggregate analysis surfaces patterns individual transcript review will never catch.

This is also why starting narrow beats starting broad. Automating everything at once produces a lower average containment rate than starting with tier 1, because the hard interactions drag the average down while the agent is still learning. Pick three high-volume, well-documented categories. Expand from evidence.

How to Set Expectations With Stakeholders

Deflection rate is the number your stakeholders ask about because it's the number vendors trained them to ask about. Reframe it early, before there's a target attached to it.

The framing that works: the goal is not to maximise how many contacts the AI handles. It's to maximise resolutions while cost per resolution falls and satisfaction holds. Deflection is one signal in that picture.

Practically, set successful containment targets at 90 days, six months, and 12 months, and say out loud that the 90-day number will be the worst one. Pair every deflection metric with a CSAT metric from the same interaction set. Track re-contact from day one. And if you're evaluating platforms, ask each vendor for their formula in writing before you compare their numbers. Otherwise you are holding a ratio next to a percentage and calling it a bake-off. That applies to the whole enterprise chatbot category, and it applies to whatever digital customer service tooling you already own.

See What Realistic Containment Looks Like for Your Operation

The right benchmark for your deployment depends on your interaction mix, your integration depth, and how your team is resourced to iterate. Voiceflow's team works with enterprise support leaders to scope containment targets against comparable customers, and to build agents designed for resolution rather than exit.

Book a personalized demo with Voiceflow →

Bring your ticket taxonomy and your current numbers. We'll tell you which of them is measuring what you think it's measuring.

Frequently asked questions

What is ticket deflection?
Ticket deflection is the share of inbound support contacts that never become a ticket handled by a human agent. The important caveat is that most conventions count an exit rather than an outcome, so a customer who abandons a chat after a useless answer is usually counted as deflected.
How do you calculate ticket deflection rate?
There is no single formula. Zendesk publishes total help-center users divided by total users in tickets, which returns a ratio such as 4. Decagon publishes deflected contacts divided by total contact attempts times 100, which returns a percentage. A third convention uses sessions, and a fourth only counts a contact as deflected if the customer did not come back about the same issue within 24 to 72 hours.
What is a good ticket deflection rate?
It depends on interaction complexity and regulatory load. Decagon publishes bands of 55% to 75% for e-commerce, 40% to 60% for SaaS, and 25% to 45% for financial services and healthcare. Gartner puts full self-service resolution at roughly 14% today and forecasts 80% autonomous resolution of common issues by 2029, so treat 80% as a destination rather than a benchmark.
What does call deflection mean?
Call deflection is the voice-channel version of the same idea: an inbound call resolved without reaching a live agent, usually by an IVR, a voice agent, or a redirect to self-service. It carries the same measurement problem, because an abandoned call and a resolved call both leave the queue.
What does case deflection mean?
Case deflection is the term used in ticketing and CRM tools such as Salesforce for a case that a customer resolves through a knowledge article or community answer before submitting it. It is measured at the point of case creation rather than across the whole contact volume.
What is the difference between deflection, containment, and resolution?
Deflection counts contacts handled without a human regardless of outcome, including abandonment. Containment counts contacts that reached the end of the AI interaction without escalating, which excludes abandonment but still does not confirm the customer got what they needed. Resolution counts issues actually solved, measured through CSAT, re-contact rate, or an explicit confirmation step. The gap between the three is the most useful diagnostic you have.
Last updated: August 28, 2026
Share this article