You already have a chatbot. That is the part every "AI agent vs chatbot" explainer skips.
Search that phrase and you get seven vendor pages telling you an agent reasons and a chatbot follows a script. All true, all beside the point if you shipped a decision-tree bot three years ago and now own the maintenance backlog, the escalation queue, and the internal reputation that came with it. The comparison is where those pages stop. It is where your actual problem starts.
So this guide does both. First the differences that change how the system behaves in production, in a table you can put in front of a stakeholder. Then the harder half: how to replace what you have without losing the coverage you already depend on, and what to expect from the result when the demo is over.
AI Agent vs Chatbot: The Short Answer
A chatbot matches input to a predefined intent and returns a predefined response. An AI agent is given a goal and a set of tools, and works out the steps itself.
Everything else follows from that one difference.
| Legacy chatbot | AI agent | |
|---|---|---|
| How it understands | Matches phrasing to a labelled intent it was trained on | Interprets what the customer is trying to accomplish, however they phrase it |
| How it acts | Returns text, links, or a menu | Calls APIs and tools: looks up the order, issues the refund, updates the address |
| Handling the unexpected | Falls back or escalates when input misses the tree | Reasons from context and available tools; asks a clarifying question |
| Conversation memory | Resets at the node; changing subject restarts the menu | Carries context across turns and topic changes |
| Cost of change | Every new policy or product means authoring new flows | Update the knowledge source or the tool; behavior follows |
| Escalation | Transfers the customer, usually with no context | Transfers a summary: what was asked, what was tried, what remains |
| How you improve it | Read transcripts, guess, edit flows, hope | Score conversations against criteria, find failure patterns, change the instruction |
The row that decides most migrations is the second one. A chatbot that answers "what is your return policy" and an agent that starts the return are not the same product with different accuracy. They sit on opposite sides of the line between deflecting a contact and resolving it, and only one of those is worth paying for.
Why Legacy Chatbots Stop Working at Scale
Rule-based bots are not badly built. They are correctly built for a scope that stops holding, which is why so much customer service automation written before 2023 is being replaced rather than extended.
Every path has to be authored in advance, every edge case anticipated, every new product or policy mapped into the tree by hand. That is fine when the scope is narrow and stable. Both conditions expire.
Four failure modes show up in roughly this order:
- Maintenance compounds. Each change to product, pricing, or policy needs a matching change to the flows. Flows accumulate faster than anyone retires them, dependencies cross, and eventually the honest answer to "what happens if a customer says X" is that nobody knows without testing it. The system is not broken. It is unowned.
- Intent matching misses. Synonyms, typos, multi-part questions, phrasing that depends on something said two turns ago. Each miss produces a fallback or an unnecessary escalation. Customers adapt by talking to the bot in stilted keyword-speak, which narrows what it ever sees, which makes the next round of training data worse.
- It answers but cannot act. This is the structural one. A bot with no path into your order system can describe a policy and nothing more. The customer still has to do the thing, or escalate to someone who can.
- Escalations arrive empty. When it does hand off, it usually hands off a transcript and a queue position. The agent restarts the conversation, the customer repeats themselves, and the interaction costs more than if the bot had never intercepted it.
That last one is worth sitting with. Gartner surveyed 5,728 customers and found that 45% of those who started in self-service said the company did not understand what they were trying to do. That is not a UX complaint. It is what an unowned decision tree feels like from the outside.
The Number That Should Set Your Expectations
Before the migration plan, the honest baseline.
In that same Gartner research, only 14% of customer service issues are fully resolved in self-service today. Gartner separately forecasts that agentic AI will autonomously resolve 80% of common issues by 2029. Those two numbers get quoted in the same breath constantly, and only one of them describes the present.
Published containment bands land between them and vary by how much regulation you carry. Decagon publishes 55% to 75% for e-commerce, 40% to 60% for SaaS, and 25% to 45% for financial services and healthcare. Plan against your band, not against the forecast, and decide which metrics you are holding the project to before the first sprint rather than after the first review.
The other thing to fix before you plan is which number you are even tracking. Vendors report deflection, and the two most-cited deflection formulas do not return the same kind of value. One gives you a ratio, the other a percentage. Deflection counts an exit. Resolution counts an outcome. A migration measured on the first can look like a win while the volume you thought you removed comes back through another channel.
Klarna is the public version of this and it went both ways. In February 2024 the company announced an assistant handling two-thirds of service chats, described as the work of 700 agents. Our teardown of that agent found replies three times longer than they needed to be and several that were simply wrong. On 8 May 2025 CEO Sebastian Siemiatkowski told Bloomberg the company had gone too far: "We focused too much on cost. The result was lower quality." Klarna resumed hiring human agents.
Nothing about the two-thirds figure was false. The volume was real and the outcome underneath it was not. Read that as the argument for measuring resolution rather than deflection, not as an argument against migrating.
How to Replace a Chatbot Without Losing the Coverage You Have
The most common way these projects go wrong is a lift-and-shift. Teams export the existing flows, rebuild them in the new platform, and end up with an agent that behaves like a marginally better chatbot, because it was specified as one.
Start from customer outcomes instead. Four questions, in this order.
1. What are customers actually trying to accomplish? Not which intents the current bot supports. Those were chosen partly for what was cheap to build in 2022. The replacement is the moment to close the gap between what you automated and what people came for.
2. Where does the current bot fail most? You already have this data and it is better than any vendor's benchmark. Pull the transcripts where the bot escalated. What was being asked, and what could it not do? Those are your highest-value automation targets, and in most support operations they cluster in the same Tier 1 interaction types: order status, password resets, billing questions, returns. They should be the first things the agent is designed to handle rather than the last.
3. What does resolution actually require? For each interaction type, trace what a human does to close it. Which systems do they open, what do they look up, what do they change. That list is your integration scope. An agent without access to those systems will deflect exactly as well as the bot you are replacing and resolve no better.
If the bot you are replacing serves employees rather than customers, the same four questions apply, but the volume tends to concentrate harder: internal help desks answer a smaller set of questions far more often.
4. What should a good escalation look like? Design the handoff before the agent. What context transfers, what triggers the transfer, what the human sees on pickup. Getting the handoff right improves the interactions the AI cannot finish, which is where most of the remaining cost sits.
Then cut over in stages rather than at once. Build and test in a development environment, promote to staging with real traffic shapes, and run the agent in production alongside the existing bot on a narrow slice of volume before widening it. Environments exist for exactly this: the old system keeps serving while the new one earns its coverage, and rollback is a promotion in reverse rather than an incident. Teams that skip the staged path are usually the ones who end up defending a regression in a leadership meeting. This is the same discipline any enterprise automation programme needs once a system starts making decisions rather than repeating steps.
Shadow mode is worth the two weeks it costs. Run the agent against live conversations without letting it reply, score what it would have said, and fix the failures before a customer meets them.
What to Look for in a Replacement Platform
Not every agent platform is built for a migration that has to preserve existing coverage. The criteria that separate them:
- Integration depth. The agent is worth exactly what it can reach. Look for connectors to your helpdesk, CRM, and order systems, plus an API and SDK flexible enough for whatever is bespoke in your stack. This is the criterion most often deferred and most often fatal.
- Deterministic paths where you need them. Refunds, cancellations, identity checks and anything a regulator reads should follow a fixed workflow, not a reasoned improvisation. The platform needs to support both: scripted workflows for the high-stakes paths, open-ended reasoning everywhere else.
- Grounding you control. Answers should come from your documentation, policies, and product data through retrieval, not from whatever the model absorbed in pretraining. You should be able to see which source produced an answer. Improving the agent then means fixing the source, not retraining a classifier.
- Conversation-level visibility. Aggregate dashboards tell you something moved. Finding out why needs individual transcripts, escalation patterns, and the ability to score conversations against criteria you define. Without that, iteration is guesswork with extra steps.
- No model lock-in. Model quality, pricing, and latency have all moved more than once a year since 2023. A platform tied to one provider means every one of those changes is someone else's decision.
- Multi-team workflows. These projects involve CX, product, engineering, and often legal. A developer-only tool pushes every content change through an engineering queue, which reproduces the maintenance drag you left the chatbot to escape. Decide who owns the agent day to day before you buy, not after.
- Security posture that survives review. SOC 2 Type 2, PII handling, data residency, and increasingly the disclosure rules that came with the EU AI Act. Work through the review criteria early. It is the stage where migrations stall.
Where Voiceflow Fits
Voiceflow is the layer above your existing stack rather than a replacement for it, which is the shape most enterprise AI deployments end up in once the integration work is scoped honestly. Agents are built on a visual canvas that CX and engineering can work in together, grounded in a Knowledge Base built from your own documentation and product data, and connected through API and SDK actions to the systems where resolution actually happens.
For the parts that must be exact, Workflows pin the steps. For everything else, Playbooks let the agent reason toward the outcome. The platform is model-agnostic across OpenAI, Anthropic and Google, so a change in model economics is a configuration change. Environments carry an agent from development to staging to production, Evaluations score conversations against criteria you define rather than a sampled handful, and Observability surfaces the failure patterns worth fixing next. Voiceflow is SOC 2 Type 2 certified with PII masking available.
Turo, StubHub, Sanlam and Trilogy run customer-facing agents on it, most of them alongside systems that were already in place.
Ready to Scope the Replacement?
A demo here is not a product walkthrough. It is a conversation about the bot you have, where it is failing, and what a staged replacement would look like against your stack and your escalation data.
Bring the escalation transcripts. They are the most useful thing in the room.