General industry terms, defined as they are used in these docs. They are not Voiceflow-specific, and are here so a page can link to one rather than explain it again. One section per term.
Models
Model
The language model generating a reply. Which one an agent uses is a behaviour setting, and the choice affects both quality and credit cost.
Also asked as:
- which AI is behind the assistant
- pick between the available LLMs
- the AI engine behind the replies
- swap the model for a cheaper one
Related: LLM · Default model · Credits
Documented in: Behaviour › Default model
LLM
Large language model. Used interchangeably with model above.
Also asked as:
- the large language model term
- what is a large language model
- the LLM the agent runs on
Related: Model
Documented in: Behaviour › Default model
Token
The unit a model reads and writes text in, roughly a word-piece. Usage and cost are measured in tokens, which is why analytics reports them.
Also asked as:
- what a token means for cost
- how text is counted by the model
- why is usage measured in tokens
- tokens versus words
Related: Context window · Credits
Documented in: Analytics
Context window
How much text a model can consider at once, counted in tokens. Everything the agent sends on a turn competes for the same window.
Also asked as:
- how much the model can keep in mind at once
- why a long prompt crowds out the conversation
- the limit on how much the model can read at once
- running out of context
Documented in: Behaviour › Memory · Global prompt › Keep it short
Temperature
How much randomness a model uses when choosing words. Lower is more repeatable, higher more varied.
Also asked as:
- make the assistant more predictable
- the creativity setting
- make replies less random
- the setting for more creative answers
Related: Model · Default model
Documented in: Behaviour › Default model · Playbooks › Model settings
Prompts
Prompt
The text given to a model to produce a reply. In Voiceflow this is assembled from the global prompt, the instructions, and the conversation so far.
Also asked as:
- what the model actually receives each turn
- everything the model is given before it answers
- the text sent to the AI each turn
Related: System prompt · Global prompt
Documented in: Global prompt › What is the global prompt?
System prompt
The instruction layer a model treats as standing orders rather than user input. Voiceflow’s global prompt fills this role.
Also asked as:
- where the standing instructions go
- the system message in Voiceflow terms
- the instructions the model always follows
- the equivalent of a system message
Related: Prompt · Global prompt
Documented in: Global prompt › What is the global prompt?
Tool calling
A model choosing to invoke a tool, with arguments, instead of replying in prose. What makes a playbook agentic.
Also asked as:
- the model decides to use a tool on its own
- function calling
- the model decides to call an API by itself
- how the agent picks a tool
Documented in: Playbooks › Tools
Retrieval
RAG
Retrieval-augmented generation. Fetching relevant documents and giving them to the model so the answer is grounded in your content rather than the model’s training. This is what the knowledge base does.
Also asked as:
- answer from our documents instead of the model’s memory
- retrieval before generation
- retrieval augmented generation in Voiceflow
- how the assistant uses our documents to answer
Related: Knowledge base · Grounding · Embedding
Documented in: Querying the knowledge base
Embedding
A numeric representation of a piece of text, used to find passages that mean something similar rather than merely sharing words.
Also asked as:
- how the knowledge base finds similar passages
- semantic search under the hood
- vector representation of text for search
- why similar wording is found even without exact words
Documented in: Importing data sources › LLM chunking strategies
Chunk
One passage a document is split into before it is embedded. Chunking decides what retrieval can return.
Also asked as:
- the passage size the bot retrieves
- how documents get split up
- a passage of a document used in retrieval
- the piece of text the bot gets back
Related: Chunking strategy · Chunk limit
Documented in: Importing data sources › LLM chunking strategies · Querying the knowledge base › Chunk limit
Grounding
Constraining an answer to supplied source material, so the model reports rather than invents.
Also asked as:
- keep the assistant to the facts in our content
- stop answers drifting from the sources
- make the assistant stick to the source material
- answers based only on provided content
Related: Hallucination · Preventing hallucination · RAG
Documented in: Global prompt › #Guardrails · Querying the knowledge base
Hallucination
A fluent, confident answer that is not true. The reason grounding and evaluations matter.
Also asked as:
- the bot makes things up
- confident but wrong replies
- the assistant states something false with confidence
- made-up answers from the model
Related: Grounding · Preventing hallucination · Evaluation
Documented in: Global prompt › #Guardrails
Protocols
MCP
Model Context Protocol. An open standard for exposing tools to a model, so the same tool works across clients.
Also asked as:
- what MCP stands for
- the standard behind tool servers
- the Model Context Protocol explained
- the standard that lets tools work across AI clients
Related: MCP server · Voiceflow MCP server · MCP tool
Documented in: Voiceflow MCP overview · MCP tool
Capitalisation carries meaning here. “Agent step” and “Function step” name specific steps; the bare words step and tool stay lowercase.