A comprehensive guide for understanding the differences between Voiceflow Legacy (V3) and V4, building agents the new agentic way, and migrating existing legacy agents to V4.
1. Overview: What Changed in V4
Voiceflow V4 represents a fundamental shift from flow-based chatbot building to agentic AI agent design. Instead of manually wiring every conversation path with rigid step-by-step logic, V4 introduces an instruction-driven, LLM-powered architecture where agents autonomously route conversations, use tools, and follow high-level instructions.
The Core Philosophy Change
Key Architectural Concepts in V4
V4 introduces a layered architecture:
- Agent Level - Global prompt (persona, tone, guardrails) + Instructions (routing logic) + System tools
- Skills Level - Playbooks (agentic, instruction-driven) and Workflows (deterministic, step-by-step)
- Tools Level - API tools, Function tools, MCP tools, Integration tools (Salesforce, Zendesk, Gmail, Make), all scoped to skills
- Knowledge Base - RAG-powered data layer with metadata filtering, integrated as a system tool
2. Deprecated Steps & Features from Legacy
2.1 Officially Deprecated Steps
AI Response Step (Deprecated)
- What it did: Generated a single AI-powered text response using a prompt. Users wrote a prompt, and the LLM produced a message.
- Why deprecated: The Message step now has a built-in Prompt mode that does exactly this - you can write a prompt to generate a single AI message. For multi-turn conversational interactions, the Playbook step is the recommended replacement.
- Migration path: Replace with a Message step (Prompt mode) for single AI-generated messages, or a Playbook step for multi-turn AI conversations.
AI Set Step (Deprecated)
- What it did: Used an LLM to extract or compute a variable value from conversation context. You would provide a prompt, and the model would fill a variable.
- Why deprecated: The Operator step and Playbook step can evaluate context and set variables using LLM reasoning. For structured data extraction, Structured Outputs with JSON schemas on tools and playbooks provide a more reliable, typed approach.
- Migration path: Replace with the Operator step (for LLM-powered conditional routing based on extracted info), Playbook step with tool outputs, or Function tools with structured outputs.
Important: Existing projects using deprecated steps will continue to work, but these steps are no longer available for new projects and will eventually lose support. Plan to migrate away from them.
2.2 Legacy Concepts No Longer Central
NLU-Based Intent Classification
- What it was: In legacy, you defined Intents with training utterances and Entities (slots) to classify user input. The NLU model matched user messages to intents for routing.
- What replaced it: V4 uses RAG-based intent recognition combined with LLM understanding. Instead of training utterances, you write Instructions in natural language that tell the agent when to route to which skill. The LLM reads the instructions, skill names, and triggers to make routing decisions.
- Migration impact: You no longer need to maintain lists of training utterances. Instead, write clear, descriptive skill names and instruction rules.
Flows / Topics / Components (Legacy Organization)
- What they were: In V3, you organized your agent into Flows (conversation paths), Topics (groupings of flows), and Components (reusable sub-flows).
- What replaced them: V4 uses Playbooks (agentic skills) and Workflows (deterministic skills) as the primary organizational units. Workflows can call other workflows using the Workflow step, and playbooks can be composed using the Crew step.
- Migration impact: Each legacy flow/topic typically maps to either a Playbook or a Workflow in V4, depending on whether the logic is better served by AI-driven or deterministic execution.
Explicit Intent-Based Triggers
- What they were: Flows were triggered by specific intents. You’d wire “if intent = order_status, go to Order Status flow.”
- What replaced them: The Agent Instructions layer handles routing. You write rules like: “If the user asks about their order status, use the Order Status skill.” The LLM evaluates these instructions on every turn.
Choice Step (Legacy)
- What it was: Presented options to users and routed based on intent matching of their response.
- What replaced it: The Buttons step (for deterministic choices) or Playbook step (for AI-interpreted user choices). The Operator step can also route based on LLM evaluation of user input.
3. New Steps & Features in V4
3.1 New Steps
Playbook Step
- Purpose: Delegates a task to an AI-powered playbook. The playbook responds dynamically to users based on instructions rather than following a fixed sequence of steps.
- When to use: When you need the AI to handle a goal-oriented task flexibly - like answering questions about a product, collecting information conversationally, or guiding a user through a process where the exact path may vary.
- Key features: Instruction-driven, can use tools (API, Function, MCP, Integrations), has its own model settings and temperature, supports structured outputs.
Crew Step
- Purpose: Groups multiple playbooks together so they can hand off tasks between each other directly. Routing happens within the step itself rather than through a separate routing agent, reducing latency.
- When to use: Customer support scenarios where you have separate playbooks for sales, billing, legal, and general inquiries, each handing off to the next as needed.
- Key features: Root playbook + subagents, direct inter-playbook handoff, lower latency than routing through the agent.
Operator Step
- Purpose: A silent, LLM-powered step that evaluates conditions and routes your workflow accordingly. Unlike code-based conditionals, it uses an LLM to interpret conditions.
- When to use: When you need intelligent routing that can evaluate variable values, user input, and session state without requiring exact string matches or rigid logic.
- Key features: No visible output to users, LLM-powered evaluation, can use tools for condition evaluation, multiple named exit paths.
MCP Step
- Purpose: Executes an MCP (Model Context Protocol) tool in your workflow. MCP is a standardized way for AI agents to interact with third-party services.
- When to use: When you need to connect to external services via MCP servers - a growing standard for AI tool interoperability.
- Key features: Input variables, capture response to variables, supports structured data responses.
Integration Step
- Purpose: Calls any third-party integration tool (Salesforce, Zendesk, Gmail, Make, etc.) directly in a workflow.
- When to use: When you need deterministic, step-by-step execution of a third-party integration at a specific point in a workflow.
- Key features: Pre-built connectors, input variable mapping, response capture.
API Step (Enhanced)
- Purpose: Executes an API tool in your workflow for calling external APIs.
- When to use: For deterministic API calls in workflows - retrieving customer data, checking inventory, processing payments.
- Key features: Input variable mapping, response capture with variable pathing for structured data, async execution support.
Function Step (Enhanced)
- Purpose: Executes a Function tool (custom JavaScript) at a specific point in your workflow.
- When to use: When you need custom logic that goes beyond built-in steps - data transformation, calculations, complex conditional logic.
- Key features: Reusable function tools, input/output variable mapping, async execution support.
Listen Step
- Purpose: Captures the user’s response and saves it to a variable. Enables two-way conversations inside workflows.
- When to use: When you need to capture freeform text input in a deterministic workflow (names, descriptions, feedback). For most conversational input, a Playbook step is recommended instead.
- Key features: Saves input to variable, enables mid-workflow user input.
Call Forward Step
- Purpose: Forwards a phone call to a specified number (voice agents only).
- When to use: When the agent needs to transfer a live call to a human agent or another department.
3.2 New Architectural Features
Global Prompt
- What: The always-on layer that shapes how your agent thinks, responds, and behaves across every turn. Defines persona, tone, and guardrails.
- Best structure: Use sections like
# Role,# Tone,# Guardrailsto organize. - Key principle: Keep it focused on identity and behavior - not task-specific logic.
Agent Instructions
- What: The decision-making layer that controls routing. Instructions tell the agent how to decide which skill to use, when to escalate, and how to handle edge cases.
- How they work: Evaluated on every turn. The LLM reads your instructions alongside skill names, triggers, and conversation context to make routing decisions.
- Key principle: Write rules, not code. Think of instructions as a human manager briefing a new employee.
Framework Selection
V4 offers three frameworks:
- Agentic (Recommended): The agent uses LLM-powered routing to decide which skill to invoke based on the user’s input, your instructions, and conversation context. Supports both agentic playbooks and deterministic workflows within the same agent.
- Conversation Flow: You design the conversation path in the visual builder step by step. No autonomous decision-making unless you explicitly add it. Works for simple, predictable use cases but becomes hard to maintain at scale.
- Hybrid (via Agentic): The agentic framework supports fully deterministic workflows - the difference is that routing between them is handled by the AI. This gives you deterministic control where you need it plus AI-powered routing that scales.
System Tools
Native capabilities the agent can use automatically when enabled:
- Knowledge Base - Search your uploaded data (RAG)
- Buttons - Present interactive button options
- Cards - Show rich cards with titles, descriptions, images, and action buttons
- Carousels - Display multiple cards in a scrollable carousel
- Call Forward - Transfer phone calls
- End Conversation - Gracefully end the interaction
Tools Architecture
V4 has a modular tools system scoped to skills:
- API Tools - Call external REST APIs
- Function Tools - Run custom JavaScript
- MCP Tools - Connect via Model Context Protocol
- Integration Tools - Pre-built connectors (Salesforce, Zendesk, Gmail, Make)
- System Tools - Built-in agent capabilities
Design principle: Tools are scoped to skills, not the agent level. This keeps prompts short, focused, and maintainable at scale.
Structured Outputs
- Define JSON schemas for tool and playbook outputs
- Ensures the LLM returns data in a predictable, typed format
- Enables reliable variable extraction without the old AI Set step
Variable Pathing
- Access nested values in structured data using dot notation
- Example:
response.customer.nameto extract a nested field from an API response
Initialization Workflow
- A special workflow that runs before the agent starts handling conversations
- Use it for: setting up variables from custom data, running API calls to pre-populate context, displaying a static start message
- Replaces legacy “launch” or “start” event handling
Events
- Trigger agent behavior from outside the normal conversation flow
- Types: Run workflow (triggers a specific workflow), Set variable (updates a value mid-conversation)
Knowledge Base Enhancements
- Metadata filtering - Add metadata tags to documents and filter which chunks are searched based on variables or context
- Table uploads - Upload CSVs where each row becomes a separate chunk with column headers as field names
- URL indexing - Index content directly from web pages
- Triggers - Add a trigger to the knowledge base tool for better retrieval accuracy
Evaluations & Testing
- Evaluations - Score agent conversations at scale to measure quality and performance
- Test Suites - Create and run automated test suites via the CLI or API
- Agent-to-Agent Testing - Use a separate Voiceflow agent (or OpenAI-powered agent) to simulate realistic multi-turn conversations and validate behavior
- Transcript-to-Test Conversion - Convert real conversation transcripts into reusable test cases
4. Building Agents in V4 vs V3: A Paradigm Shift
V3 Approach: Flow-Based Design
In V3, building an agent meant:
- Define intents and entities - Create a list of intents with training utterances and entity slots for NLU classification.
- Build flows - Create visual flow diagrams where each intent triggered a specific flow.
- Wire every path - Manually connect every step to the next, handling every branch, fallback, and edge case with explicit connections.
- Sprinkle AI - Add AI Response or AI Set steps where you needed LLM-generated content or extraction.
- Handle fallbacks manually - Create “no match” and “no reply” handlers for each flow.
The result: A web of interconnected flows that became increasingly fragile as complexity grew. Adding a new feature meant rewiring exit conditions across multiple flows. Edge cases multiplied, and touching one path could break another.
V4 Approach: Agentic Design
In V4, building an agent means:
- Define the agent’s identity - Write a Global Prompt that sets persona, tone, and guardrails. This applies to every turn.
- Write routing instructions - Define Instructions that tell the agent which skill to use and when. Rules, not wires.
- Create skills - Build Playbooks (for AI-driven tasks) and Workflows (for deterministic logic). Each skill is self-contained with its own instructions and tools.
- Attach tools to skills - Scope API calls, functions, integrations, and MCP connections to individual skills. Each skill owns its own tools.
- Enable system tools - Turn on knowledge base, buttons, cards, etc. at the agent level. The agent decides when to use them.
- Test and iterate - Use evaluations, test suites, and transcripts to measure quality and refine.
The result: A modular, scalable agent where adding a new capability means creating a new skill and adding an instruction line - no rewiring required.
Side-by-Side Comparison
5. Best Practices for Designing Agents in V4
5.1 Agent-Level Design
Keep the Global Prompt Focused
- Do: Define persona, tone, language, and guardrails.
- Don’t: Put task-specific logic or tool instructions in the global prompt.
- Why: At enterprise scale, a bloated global prompt with dozens of tool instructions becomes fragile and slow. Keep it under 500 words.
- Structure it clearly: Use sections like
# Role,# Tone,# Guardrails.
Write Clear, Specific Instructions
- Instructions are the decision-making layer. Write them like a manager briefing a new employee.
- Be explicit about routing: “If the user asks about X, use the Y skill.”
- Include edge cases: “If the user’s request is ambiguous between billing and technical support, ask a clarifying question.”
- Include escalation rules: “If the user asks to speak to a human, use the Escalation skill.”
Name Skills Descriptively
- Skill names are visible to the LLM during routing. A clear name helps the model choose correctly.
- Good: “Check Order Status”, “Process Refund Request”, “Book Demo Appointment”
- Bad: “Flow 1”, “Handler”, “Misc”
Write Good Triggers
- Each skill should have a one-sentence, outcome-focused trigger.
- Example: “Helps the user check the current status of an existing order using their order ID.”
- This is what the LLM uses (alongside instructions) to decide whether to invoke the skill.
5.2 Skill Design
Choose Playbook vs Workflow Wisely
Use a Playbook when:
- Flexibility matters more than rigid control
- The task involves natural conversation (Q&A, information gathering, recommendations)
- You want the AI to decide which tools to use and when
- The exact conversation path may vary
Use a Workflow when:
- You need guaranteed execution order (eg: “always check inventory before confirming order”)
- The logic involves precise branching, conditions, and data transformations
- You need deterministic behavior with no AI interpretation
- Compliance or auditability requires predictable steps
Scope Tools to Skills
- Attach tools (API, Function, MCP, Integrations) to the skill that uses them - not at the agent level.
- This keeps each skill self-contained and prevents prompt bloat.
- If multiple skills need the same API, create the tool once and attach it to each skill.
Use the Crew Step for Multi-Agent Workflows
- When you have multiple related playbooks that need to hand off between each other, use a Crew step instead of routing through the agent.
- This reduces latency and provides smoother handoffs.
5.3 Data & Variables
Use Structured Outputs
- When extracting data from conversations or APIs, define JSON schemas for structured outputs.
- This gives you reliable, typed data instead of unpredictable LLM text.
Use Variable Pathing
- Access nested data with dot notation:
api_response.data.customer.email - Reduces the need for Code steps to parse JSON.
Use the Initialization Workflow for Setup
- Pre-populate variables, run setup API calls, or display a static welcome message before the agent starts.
- Pass custom variables from the widget and process them here.
5.4 Knowledge Base
Structure Data for Retrieval
- Use metadata tags to enable filtered searches.
- Upload CSVs as tables - each row becomes a chunk, column headers become field names.
- Add a trigger to the knowledge base tool to improve retrieval accuracy.
Use Metadata Filtering
- Tag documents with categories, product lines, user tiers, etc.
- Configure the knowledge base tool to filter based on variables - eg: only search billing docs when the user is in the billing skill.
5.5 Testing & Quality
Use Evaluations
- Define evaluation criteria and score conversations at scale.
- Run evaluations on batches of transcripts to identify quality issues.
Build Test Suites
- Create automated test suites that validate expected agent behavior.
- Convert real transcripts into test cases for regression testing.
- Use agent-to-agent testing for realistic multi-turn conversation validation.
6. Best Practices for Migrating Legacy Agents to V4
6.1 General Migration Philosophy
Do not try to replicate your V3 agent 1:1 in V4. The paradigm is fundamentally different. Instead, re-think your agent from the ground up using V4’s agentic architecture, using your V3 agent as a reference for what the agent needs to do - not how it should do it.
6.2 Step-by-Step Migration Strategy
Phase 1: Audit Your Legacy Agent
- List all intents - Document every intent, its training utterances, and what flow it triggers.
- List all flows/topics - Document the purpose of each flow, what it does, and what data it needs.
- List all API calls - Document every external API call, what data it sends/receives, and where it’s used.
- List all entities/variables - Document every entity and variable, what data it holds, and where it’s used.
- Identify AI steps - Find all AI Response and AI Set steps and document their prompts.
- Map conversation paths - Document the main happy paths and edge cases your agent handles.
Phase 2: Redesign for V4
- Define the agent identity - Write a Global Prompt based on your legacy agent’s persona and behavior.
- Group flows into skills - Each legacy flow/topic typically becomes a skill:
- Flows with flexible, conversational logic → Playbooks
- Flows with strict, sequential logic → Workflows
- Simple FAQ/knowledge flows → May not need a skill at all (just the Knowledge Base system tool)
- Write Instructions - Convert your intent-to-flow mapping into instruction rules:
- Legacy: Intent “order_status” → Order Status Flow
- V4: “If the user asks about their order status or tracking, use the Order Status skill.”
- Convert API calls to Tools - Create API Tools for each external API and scope them to the relevant skills.
- Convert AI steps - - AI Response → Message step (Prompt mode) or Playbook instructions
- AI Set → Structured outputs on tools/playbooks, or Operator step
- Set up Knowledge Base - Migrate any FAQ data, documentation, or reference material into the Knowledge Base with proper metadata.
Phase 3: Build and Test
- Start with the agent level - Set up Global Prompt, Instructions, and System Tools.
- Build skills one at a time - Start with the most critical skill, test it, then move to the next.
- Test routing first - Before perfecting individual skills, verify that the agent routes to the correct skill for various inputs.
- Use real transcripts - Pull transcripts from your V3 agent and use them to test V4 behavior.
- Run evaluations - Set up evaluations to measure quality against your legacy agent’s performance.
6.3 Common Migration Patterns
Pattern: Intent-Triggered FAQ Flow → Knowledge Base + Instructions
Legacy (V3):
V4:
No skill needed - the agent handles it directly with the knowledge base.
Pattern: Data Collection Flow → Playbook with Structured Output
Legacy (V3):
V4:
The playbook handles the conversation naturally - it may collect all info in one message or across multiple turns depending on what the user provides.
Pattern: Conditional Routing Flow → Workflow with Operator
Legacy (V3):
V4:
Or, even simpler - let the Agent Instructions handle this routing entirely without a workflow:
Pattern: Complex API Integration Flow → Workflow with API/Function Steps
Legacy (V3):
V4:
This stays deterministic because refund processing requires guaranteed execution order.
7. Step-by-Step Migration Checklist
Use this checklist to track your migration progress:
Pre-Migration
- [ ] Document all legacy intents and their training utterances
- [ ] Document all legacy flows/topics and their purposes
- [ ] Document all API integrations and their endpoints
- [ ] Document all entities/variables and their usage
- [ ] Identify all AI Response and AI Set steps
- [ ] Export legacy agent transcripts for testing
Agent Setup
- [ ] Create a new V4 project
- [ ] Write the Global Prompt (persona, tone, guardrails)
- [ ] Choose the Agentic framework (recommended)
- [ ] Configure default model and temperature in Behaviour settings
- [ ] Set up the Initialization Workflow (if needed)
Knowledge Base
- [ ] Upload all FAQ/documentation data
- [ ] Add metadata tags for filtering
- [ ] Configure knowledge base system tool with trigger
- [ ] Test knowledge base retrieval accuracy
Skills
- [ ] Map each legacy flow to a Playbook or Workflow
- [ ] Create each skill with clear name and trigger
- [ ] Write instructions for each Playbook
- [ ] Build step logic for each Workflow
- [ ] Create and attach API Tools to relevant skills
- [ ] Create and attach Function Tools to relevant skills
- [ ] Set up Integration Tools (Salesforce, Zendesk, etc.) if applicable
Routing & Instructions
- [ ] Write Agent Instructions for skill routing
- [ ] Include edge case handling in instructions
- [ ] Include escalation rules
- [ ] Add a start message (via instructions or initialization workflow)
System Tools & Behaviour
- [ ] Enable/configure Knowledge Base system tool
- [ ] Enable/configure Buttons system tool (if needed)
- [ ] Enable/configure Cards and Carousels (if needed)
- [ ] Configure global no-match behaviour
- [ ] Configure global no-reply behaviour
- [ ] Set up Events (if using external triggers)
Testing
- [ ] Test routing with diverse user inputs
- [ ] Test each skill individually
- [ ] Test edge cases and fallback behaviour
- [ ] Run legacy transcripts through V4 agent
- [ ] Set up and run evaluations
- [ ] Create automated test suites
- [ ] Compare V4 quality metrics against V3 baseline
Deployment
- [ ] Publish V4 agent
- [ ] Configure chat widget (or phone/API deployment)
- [ ] Monitor initial transcripts closely
- [ ] Iterate on instructions and skills based on real conversations
8. Reference: V3 to V4 Concept Mapping
9. FAQ
Can I use both Playbooks and Workflows in the same agent?
Yes. The Agentic framework supports both. Use Playbooks for flexible, AI-driven tasks and Workflows for deterministic, step-by-step logic. The agent’s Instructions handle routing between them.
Do I need to retrain intents in V4?
No. V4 uses RAG-based intent recognition. Instead of training utterances, you write clear skill names, triggers, and instruction rules. The LLM handles understanding.
Can I still build fully deterministic agents in V4?
Yes. You can use the Conversation Flow framework for fully deterministic agents, or use the Agentic framework with Workflows for deterministic execution within an AI-routed agent. You can also mix - use Workflows for tasks that need guaranteed execution and Playbooks for everything else.
What happens to my existing V3 projects?
Legacy projects continue to work. Deprecated steps (AI Response, AI Set) remain functional in existing projects but are no longer available for new projects. Plan to migrate before support is eventually removed.
Can I export my V3 agent and import it into V4?
There is no automatic migration tool. You should create a new V4 project and rebuild using V4’s architecture, using your V3 agent as a reference. The CLI can export agent data (voiceflow-cli agent export) which can help document your V3 agent’s structure.
How do I handle multi-language agents in V4?
Set the language in your Global Prompt and Behaviour settings. The LLM naturally handles multi-language conversations. You can also use instructions to define language-specific behavior (eg: “If the user writes in French, respond in French”).
What if I need an API call at the agent level, not scoped to a skill?
This is intentionally not supported. Voiceflow recommends keeping the agent level focused on identity and routing. Create a skill for the API call and route to it. This keeps your architecture modular and maintainable.
How do I replicate the old “Go To” functionality?
Use the Workflow step to jump to another workflow, the Playbook step to delegate to a playbook, or simply let the Agent Instructions route to the appropriate skill.
What’s the difference between the Global Prompt and a Playbook prompt?
The Global Prompt applies to every turn of every conversation - it shapes the agent’s overall behavior, persona, and guardrails. A Playbook prompt (instructions) only applies when that specific playbook is active, defining task-specific behavior and goals.
Additional Resources
- Voiceflow V4 Documentation
- Choosing a Framework
- Global Prompt Guide
- Agent Instructions
- Playbooks
- Workflows
- Tools Overview
- System Tools
- Knowledge Base
- Deprecation Notice: AI Response & AI Set Steps
- Voiceflow CLI
Last updated: March 2026