Improved

Instruction patching

Agent instructions, global prompts, and playbook instructions can now be read, searched, and patched line by line through the API. You can edit a single line in a 20,000-line prompt without re-sending the entire document, and search instructions to locate specific lines.

Use the new patch endpoints to update individual lines by their line number or search pattern. This applies to instructions at the agent level, global prompt level, and within playbooks.

Full document replacement remains available. Line-by-line patching is an alternative for large instruction sets where targeted edits are more efficient.

Improved

Atlas reads your conversations

Atlas can now answer questions about your conversations in bulk. Ask why calls dropped last Tuesday, or which conversations scored badly on an evaluation, and it searches your transcripts, reads the ones that matter, and reports back, instead of you opening them one at a time.

Point it at a date range, an environment, or an evaluation score. It can summarize a single conversation from a link, or run named evaluations across everything it found in one batch. Batches spend your workspace’s evaluation quota and write scores onto the transcripts, so Atlas tells you the cost before it starts.

Changed

Reasoning effort

You can now set how hard a Claude model thinks before it answers. Pick a Claude model served through Bedrock and a Reasoning effort slider appears, with stops for Off, Adaptive, Low, Medium and High. Leave it alone and the model keeps its own default.

Set it in SettingsBehaviourModel & reasoning, or from the Model popover on a playbook or prompt. Thinking tokens are billed as output, so Off is a cost control as much as a speed one.

Reasoning effort
Improved

Phone call audio

We have improved how phone call audio is processed before the agent hears it. There were cases where noise cancellation was stripping parts of what the caller said, so the agent missed some of the utterance and it did not reach the transcript. Callers now come through in full.

Nothing to configure, and this applies to every phone call.

Added

Evaluation run scope

An evaluation can now be scoped to conversations that ran on a published environment, or to draft ones, instead of scoring everything. Resolution rate and Conversion rate are published-only by default, so test traffic no longer skews them.

Set runScope to published, draft or both through the API. It applies to automatic scoring only: a batch run from the Transcripts tab still runs on exactly the transcripts you pick.

Evaluations score what already happened. Tests pin what must happen.

Changed

Billing user role

Billing access is now decided by a member’s workspace role alone. Owner, Admin and Billing roles keep it, as do organization admins.

Giving someone edit or admin access to a single project used to carry billing access with it. If you were relying on that, give them the workspace Billing role instead.

Billing user role
Improved

Atlas writes your tests

Atlas can now create tests in the Tests tab: scripted turn-by-turn conversations, persona simulations, and checks on what the agent said, where it routed and which tools it called. It can also turn a conversation into a regression test, whether that conversation worked or failed.

After a smoke conversation passes, Atlas offers to pin it as a test. Before a publish or merge to Main, it offers to run the suite first. Atlas writes the test and hands the run back to you to trigger.

Improved

Atlas improvements

Atlas now edits prompts in front of you. Ask it to change your global prompt, agent instructions or a playbook’s instructions, and you watch the rewrite happen in the editor rather than waiting for a finished result. In Safe mode the proposal appears as a word-level diff you accept or reject; in Autonomous mode the text is typed straight in.

You can also ask about one part of a prompt instead of the whole thing. Select any text in an editor and choose Ask Atlas from the toolbar above the selection, and that passage is attached to your next question as a chip.

And Atlas can set up credentials with you. When something needs a secret, an integration or an MCP server, it offers a card that opens the normal Voiceflow flow. You fill it in, the card reports back, and Atlas carries on from where it stopped. The credential is entered by you and never passes through the chat or the model.

Added

Secret and integration endpoints

Secrets and third-party integrations can now be listed, created, connected and disconnected through the stable API, the CLI and the Voiceflow MCP server, scoped by project. Secret values are write-only and never returned.

Integrations report whether they can be connected through the API at all. Browser-flow ones such as Zendesk and Salesforce still have to be connected in the builder. A credential value is refused when it arrives from an MCP client.

Improved

Pending changes indicator

An environment with unpublished work now says so. The Publish control carries an amber dot, and its menu offers View pending changes so you can read the diff without opening the Publish form. The environments table shows the same dot in place of the old Draft tag.

Pending changes indicator
Improved

Safe mode in Atlas

Safe mode is how Atlas works by default: it pauses and asks before anything that changes your project, and runs read-only work such as listing, fetching, exporting, querying analytics and searching transcripts without interrupting you.

It now decides what counts as a change per operation, rather than inferring it from the operation’s name. Anything implemented as a write asks first even if it reads like a lookup, and any newly added operation asks until it is classified. Switch between Safe and Autonomous from the pill in the Atlas composer footer.

Atlas Safe mode
Added

Atlas

Meet Atlas, your in-app copilot for building and monitoring agents. Ask how an agent performed this week, how that compares to last week, or where customers are dropping off. Ask for a change to the agent and Atlas makes it.

Atlas

Atlas holds a single conversation as you move around the product, so you never have to re-explain what you are working on. You can watch it think: a live activity line shows the tools it is calling and the reasoning behind each step.

Anything you find yourself asking repeatedly can become a Skill. Save your own in Settings → Skills, then pick them from Atlas whenever you need them.

Atlas is in Beta, available on agentic projects, and its usage is broken out in your workspace usage charts.

Added

Markdown source view

The fullscreen global prompt and playbook instruction editors can now be switched to their markdown source and edited as plain text, so you can see and change exactly what the agent will read. Use Show markdown source in the floating actions, and Show rendered markdown to switch back.

Source mode has syntax highlighting and tab indentation, and your choice is remembered per editor. The rendered view is still the default, so nothing changes unless you switch. Copying from the rendered view now gives you markdown.

Markdown source view
Added

You can now find transcripts by what was actually said in them. Search for a phrase and the list narrows to conversations containing it, with each result showing how many messages matched and a short excerpt around the first hit.

You can also narrow to just what the user said, or just the agent. Search terms can be up to 256 characters, and the same filter is available on the transcript search API.

Transcript keyword search
Improved

Max conversation duration

Voice conversations can now run up to sixty minutes, double the previous ceiling. You can set the limit from Behaviour → Session & timeout.

Agents that had the setting switched off move to a thirty-minute default. If you previously entered a value above thirty minutes, note that it was silently capped before and will now apply in full.

Max conversation duration
Improved

MCP token authentication

The Voiceflow MCP server now accepts a personal access token as a bearer credential, so a client that cannot open a browser for OAuth can connect using the same token that authenticates the REST API and the CLI. Send it in the Authorization header instead of completing the browser sign-in. The OAuth flow is unchanged.

Improved

Trace filtering on the conversation endpoint

A caller of the stable conversation endpoint can now list trace types to drop from the response, so an integration only receives the traces it renders. Send a config object alongside the action with excludeTypes listing the types to drop. Omit it and every trace comes back as before.

One trace is always returned even when its type is excluded: the debug trace reporting that the credit limit was reached.

Improved

Late variable writes

A variable set after the agent has already replied, which is the normal case for an asynchronous tool that finishes late, now survives into the following turn instead of being quietly reverted. The agent reads what was actually written.

Nothing to configure. If your agent was built to expect late writes to be discarded, note that it will now see the updated values instead.

Improved

Multilingual Cartesia voices

Cartesia text-to-speech now covers over forty languages, up from fifteen. Each voice declares which languages it speaks, and the preview sample plays in the language you selected. Choose a provider, voice and language from Settings → Behaviour → Voice output.

The voice menu now filters by language, so a voice that does not support the language you pick will not appear in the list. Existing selections are untouched.

Improved

Knowledge base logs in evaluations

Knowledge base logs joins Playbook, Workflow and Tool logs in an evaluation’s Log visibility settings. With it on, the evaluating model can see the knowledge base searches your agent ran, so a criterion about whether an answer came from your content is something it can actually check. As with the other log types, including them makes each prompt longer and raises the average cost per evaluation.

Knowledge base logs
Improved

Transcript environment fields

A transcript fetched from the REST API now reports which environment the conversation ran in, and whether it ran against the draft or published version. Both fields are null for conversations recorded before environments existed.

Added

Live agent typing indicator

When a live agent is composing a reply during a handoff, the user now sees that the agent is typing instead of waiting in silence.

Nothing to configure. This currently applies to handoffs to Dixa; other providers are unchanged.

Improved

Parallel tool calls

When an agent makes several tool calls in a single turn, all of them now run and every result comes back. Previously all but the last were discarded, so an agent that searched the knowledge base three times answered from one search.

Nothing to configure. This covers every kind of tool call an agent makes in a turn, including knowledge base searches, API tools, functions and integrations. Each result is matched back to the call that asked for it, and turns that make several calls are faster.

Added

Tool message traces

The message an agent speaks while a slow tool runs now reaches applications that call the API without streaming. It arrives as an ordinary text trace for chat or a speak trace for voice, generated through the same path as any other agent message. Previously it was dropped.

Nothing to configure. If your integration consumes completion events, it keeps receiving the message as a completion sequence. If it does not, responses from non-streaming endpoints now carry one additional trace.

Improved

Out-of-credit call handling

Phone calls are now declined when your organization has run out of credits, instead of connecting to an agent that cannot respond. The caller hears that no agents are available. A test call from the builder tells you to add credits.

Organization admins are emailed when this occurs.

Use the up and down arrow keys to select a result, Enter to open it, and Escape to close the search.