The promise of AI in customer service is easy to state: fewer tickets reaching your human agents, faster resolutions, lower cost per interaction, and a support operation that scales without scaling headcount.
The business case is harder. Not because the return isn't real, but because enterprise buyers are rightly skeptical of vendor math that conveniently lands on a 400% return. Skeptical buyers deserve better than a counter-offer of the same math with smaller numbers.
So this guide does something the genre usually avoids. Every figure below comes from a published source, named inline, and where the honest answer is unflattering to us, that's the number you get. When a CX leader takes an AI investment to their CFO, what they need is a model that survives a second reading.
What Support Actually Costs Today
Most ROI conversations skip straight to savings. The CFO will start one step earlier, because savings are meaningless without a denominator.
Your fully loaded support cost is not your headcount bill. It's:
- Agent compensation, loaded: salary, benefits, payroll tax, typically 1.25 to 1.4× base
- Management and QA overhead: team leads, workforce management, quality review
- Recruiting and ramp: cost per hire, plus the weeks a new agent is paid but not yet productive
- Tooling: helpdesk licences, telephony, workforce management, knowledge management
- Facilities and infrastructure, or the remote equivalent
- The cost of peak staffing: the agents you carry through eleven quiet months to survive the twelfth
Divide that by resolved contacts and you have your real cost per contact. Most teams are surprised by it, because the tooling and ramp lines rarely sit in the support budget.
For a benchmark, Gartner's cost-per-contact research puts the median at $1.84 for self-service and $13.50 for assisted channels, with phone, chat, and email landing in a similar band as each other. That roughly 7× gap between a contact your systems handle and a contact a person handles is the entire economic engine of this category. Every model below is a way of asking how many contacts you can move across it, and what it costs to move them.
If you're building this baseline properly, the metrics that actually belong in it are worth settling before you start, because the definitions vary more than most teams expect.
The Formula, Stated Plainly
Here is the calculation. It fits in a spreadsheet and takes about ten minutes.
Annual gross savings = (monthly contacts × containment rate × (assisted cost per contact − self-service cost per contact) × 12) + (agent hours saved × loaded hourly cost)
Net ROI = (annual gross savings − annual total cost of ownership) ÷ annual total cost of ownership
Four inputs decide the answer, and only one of them is about the AI:
- Monthly contact volume. You already know this.
- Cost per contact, both channels. Use your own numbers if you have them, Gartner's medians if you don't.
- Containment rate. The percentage of contacts fully resolved without a human. This is the input everyone gets wrong, and the next section is about why.
- Total cost of ownership. Platform, implementation, integration, and the ongoing headcount to run it. Not just the licence.
Run it three times: conservative, base, optimistic. A single number reads as a pitch. A range reads as analysis, and it pre-empts the "that seems too good to be true" objection before your CFO has to raise it.
Containment Is the Input Everyone Inflates
Take a contact centre handling 10,000 contacts a month, with Gartner's medians as the channel costs. The spread between channels is $11.66 per contained contact. What changes the answer is containment:
| Containment | Contacts contained / mo | Monthly saving | Annual saving |
|---|---|---|---|
| 25% | 2,500 | $29,150 | $349,800 |
| 40% | 4,000 | $46,640 | $559,680 |
| 60% | 6,000 | $69,960 | $839,520 |
Half a million dollars separates the top row from the bottom. So the number you assume here is the business case, and it deserves more scrutiny than the rest of the model combined.
Two published anchors are worth holding in view. Gartner's survey of 5,728 customers found that only 14% of customer service issues fully resolve in self-service, and only 36% even for issues customers themselves described as very simple. That's the floor the category is climbing out of. At the other end, Decagon publishes containment bands by vertical: 55–75% for e-commerce, 40–60% for SaaS, 25–45% for financial services and healthcare. Those are mature-deployment figures, not month-three figures.
Which means the honest planning assumption is a curve, not a constant. Start near the bottom of your vertical's band, and model the climb as something you earn over quarters.
It also means the containment number a vendor quotes you is close to meaningless without its definition attached. What counts as a deflection varies wildly between vendors. Some publish a ratio, some a percentage, and the two aren't comparable. Ask how the denominator is built before you put the number in a model.
One more thing the table hides: not every contained contact is a saved contact. A conversation that ends because the customer gave up costs you more than the ticket you avoided. So pair containment with a clean handoff path to a human and measure the abandonment rate alongside it.
Where the Savings Actually Come From
Enterprise teams realise value in three buckets. The first rarely justifies the investment alone; the combination usually does.
1. Direct Cost Reduction Through Deflection
This is the table above, and in year one it's typically the largest line. Every contact resolved without a human saves the difference between your two channel costs. The work of capturing it is unglamorous: identifying which contact types are genuinely automatable, grounding the agent in content that actually answers them, and expanding coverage one category at a time. The mechanics of getting a self-service layer to resolve rather than deflect are where most of the effort goes.
2. Agent Productivity
AI doesn't only remove interactions. It makes the ones that still reach a person faster to handle.
The best evidence here is a field study rather than a vendor claim. In Generative AI at Work (Brynjolfsson, Li and Raymond, NBER, published in the Quarterly Journal of Economics), researchers tracked 5,179 customer support agents through a staggered rollout of an AI assistant. Productivity, measured as issues resolved per hour, rose 14% on average.
The distribution matters more than the average. The gain was 34% for novice and low-skilled agents, and close to nothing for experienced ones. The assistant works by spreading what the best agents already know.
For a 20-agent team, 14% is the effective equivalent of roughly 2.8 additional agents. That's a real number and a smaller one than most vendor decks imply. But it tells you something a bigger number wouldn't: this line item is worth most to teams with high attrition, heavy seasonal hiring, or a long ramp. If your agents average five years of tenure, model this bucket at close to zero. If you rehire a third of the floor every year, it may be the largest number in your case. The same logic explains why teams under headcount pressure see the steepest returns, and why staff augmentation math and automation math should be run against each other rather than in separate documents.
3. Revenue Protection
This bucket is the one most often asserted and least often evidenced, so treat it carefully.
The mechanism is sound: support failures drive churn, and faster resolution reduces failures. What's missing is a defensible public number to multiply by. Retention sensitivity to support quality varies so much by business model that borrowed benchmarks are worse than useless here.
So don't borrow one. If you have churn data tied to support outcomes, you have the only figure worth using, and it will be more persuasive than anything you could cite. If you don't, present this bucket qualitatively and say so. A CFO will respect the distinction between a modelled number and an argued one, and will notice if you blur it. Customer effort score is usually the cheapest instrument for starting to build that data.
There's a direct revenue angle too. Agents deployed for upsell, cross-sell, and lead qualification turn support contacts into revenue moments. Same caution applies: model it only if you can measure it.
The Ceiling Worth Anchoring On
If you want one external number for the top of your range, Gartner forecasts that by 2029 agentic AI will autonomously resolve 80% of common customer service issues, producing a 30% reduction in operational costs.
Read it as a destination and a ceiling. It's a three-year forecast about common issues, not a benchmark for your next budget cycle. Anchoring your optimistic case at a 30% operational cost reduction is defensible. Anchoring your base case there is the mistake this whole guide exists to prevent.
Why 40% of These Projects Get Cancelled
Here's the number that belongs in your business case even though no vendor wants to put it there.
Gartner predicts that over 40% of agentic AI projects will be cancelled by the end of 2027, citing escalating costs, unclear business value, and inadequate risk controls. The same research calls out "agent washing" as widespread: existing chatbots, assistants, and RPA tools rebranded as agents without the underlying capability. Gartner's estimate is that only around 130 of the thousands of vendors claiming agentic capability are the real thing.
Including this in your proposal does three useful things. It shows you've read the sceptical literature, not just the vendor decks. It gives you a reason to propose staged funding rather than a single large commitment. And it reframes the evaluation question from which platform has the best demo to which platform can we still be improving in eighteen months. As the next section argues, that is the question that actually determines the return.
What Separates High-ROI Deployments From Low Ones
The difference between a deployment that pays back in months and one that delivers marginal savings almost never comes down to the underlying model. It comes down to how the deployment is run.
- High-ROI teams treat the agent like a product. There's an owner, a roadmap, and a regular cadence of review. They look at what the agent got wrong last week and fix it. Low-ROI teams deploy and move on.
- High-ROI teams start narrow and expand. One product line, one channel, or one query category, optimised to high containment, then rolled outward. Low-ROI teams automate everything at once and end up with mediocre coverage everywhere. This is also the safest path through a pilot into production.
- High-ROI teams measure leading indicators. Containment and cost per resolution are lagging. Conversation completion rate, escalation rate by topic, and CSAT on AI-handled contacts tell you where to intervene before the lagging numbers move.
- High-ROI teams pick platforms built for iteration. Closed systems that hide conversation data, or require vendor involvement to change a workflow, put a ceiling on the return no amount of model quality can lift.
That last point is worth being concrete about, because "built for iteration" is exactly the kind of phrase that gets asserted and never specified. In practice it means a handful of capabilities you can check for in a demo:
Workflows and Playbooks together. Deterministic workflows for the paths that must go the same way every time (refunds, verification, compliance checks), and playbooks for the open-ended ones where the agent reasons toward a goal. A platform that offers only one of the two forces you to choose between control and coverage, and the ROI case needs both.
Knowledge-base grounding. Answers drawn from your actual content, so improving the agent is often a content task your existing team can do rather than an engineering ticket.
Model-agnostic routing. The ability to move between providers as pricing and capability change. Over a three-year TCO horizon this is a cost line, not a technical preference.
Evaluations and Observability. These produce the leading indicators in the bullet above. Without them, "measure what matters" is advice you can't act on. Voiceflow's platform exposes both, along with separate dev, staging, and production environments so changes are tested before customers see them.
Security and compliance you can hand to procurement. SOC 2 Type 2, PII masking, and data-residency answers ready before the security review starts, not after. The enterprise security and compliance checklist covers what to ask for.
Teams running this way include Turo, StubHub International, Sanlam Studios, and Trilogy, across contact-centre operations, help-desk automation, and multilingual support where the cost-per-contact gap is widest.
The Honest Answer
The return is real, and it compounds. It's also smaller in year one than most proposals claim, concentrated in places the proposals rarely name, and dependent on execution rather than model choice.
A business case that says all of that is more likely to get funded than one that promises 400%. Finance leaders approve proposals they can defend to someone else. Give them the ranges, the sources, the cancellation base rate, and the reason your deployment won't be in that 40%.
The variable was never whether the ROI exists. It's whether your operation is set up to capture it, and whether your proposal is honest enough to survive the meeting.
A good place to start is the low end of your own volume: the tier-one contacts you already know are automatable, and the broader customer service automation picture they sit inside.
Build the Model for Your Own Numbers
The figures here are benchmarks, not predictions. Your actual return depends on contact volume, channel mix, current cost structure, and the complexity of your inbound.
Voiceflow's team works with enterprise CX leaders to build a deployment model against their real numbers, including containment projections drawn from comparable customers rather than best case.
Book a demo and you'll leave with a range you can take to your CFO, with the assumptions written down next to it.