Overview
This guide covers configuring AssemblyAI’s Universal-3.6 Pro speech-to-text model as the transcriber for a Vapi voice assistant, using the Vapi API. Vapi is a hosted platform for building voice agents. You define an assistant — its transcriber, LLM, voice, and turn-taking behavior — as a JSON object, then connect it to phone numbers, web calls, or SIP. Vapi runs the real-time pipeline for you, so there’s no agent server to deploy.Universal-3.6 Pro is our flagship next-generation streaming model for voice agents — multilingual and promptable, with conversation context.Set
transcriber.provider to "assembly-ai" and transcriber.speechModel to "universal-3-6-pro". Vapi defaults to universal-streaming-english, so set the model explicitly.Turn detection
Decide when the caller is done speaking, and how Vapi uses AssemblyAI’s end-of-turn signal.
Latency
Shorten the gap between the caller finishing and the assistant replying.
Accuracy
Prompting, key terms, language steering, and conversation context.
Interruptions
Barge-in and backchannels with Vapi’s stop speaking plan.
For a standalone voice agent without Vapi, see the AssemblyAI Voice Agent API, which handles STT, LLM routing, and TTS in a single WebSocket connection.
Quickstart
Create an assistant with Universal-3.6 Pro through the API, then call it.1
Get your API keys
You need a Vapi private API key from the Vapi dashboard. Vapi’s API authenticates with it as a bearer token:An AssemblyAI API key is optional. By default, Vapi provides access to AssemblyAI through its own account and bills transcription on your Vapi invoice. To use your own AssemblyAI account instead, see the next step.
2
Connect your AssemblyAI key (optional)
To have AssemblyAI bill transcription directly, add your AssemblyAI API key to Vapi. In the dashboard, open Integrations and add AssemblyAI, or create the credential through the API:Once the key is validated, AssemblyAI transcription runs on your AssemblyAI account and Vapi no longer charges for it. See Vapi’s provider keys guide.
3
Create an assistant
Create an assistant with a Universal-3.6 Pro transcriber. The request leaves The response includes the assistant’s
startSpeakingPlan.smartEndpointingPlan unset, which is what tells Vapi to use AssemblyAI’s end-of-turn detection:- cURL
- Python
- JavaScript
id. Always set voice explicitly: if you leave it out, Vapi assigns a default voice rather than the one you expect. The example uses Vapi’s built-in Elliot voice; see Vapi’s Create assistant reference for every voice provider and field.minEndOfTurnSilenceWhenConfident and maxTurnSilence above are the values the balanced preset uses, so they behave the same as mode alone; keeping them in the request makes the timing explicit and gives you the two values to tune. vadAssistedEndpointingEnabled: false makes Vapi act on AssemblyAI’s end-of-turn as soon as it arrives. See Turn detection for the recommended settings.This guide calls the REST API directly. Vapi’s server SDKs (
@vapi-ai/server-sdk, vapi_server_sdk) wrap the same endpoints, but their generated types can lag behind the API. If your SDK version’s types don’t include universal-3-6-pro yet, upgrade the SDK or send the request with the REST API.4
Call the assistant
The quickest test is a web call from the dashboard: open Assistants, select your assistant, and select Talk.To have the assistant call a phone, pass its To answer inbound calls, assign the assistant to a phone number in the dashboard. To embed the assistant in a web page, use the Vapi Web SDK. After each call, the call log in the dashboard shows the AssemblyAI transcript.
id and a Vapi phone number to POST /call:Switching an existing assistant to Universal-3.6 Pro
To move an existing assistant to AssemblyAI, update itstranscriber with PATCH /assistant/{id}:
startSpeakingPlan.smartEndpointingPlan, remove it too, so Vapi uses AssemblyAI’s end-of-turn detection instead of its own. See Turn detection.
In the dashboard, the same settings are on the assistant’s Transcriber tab: set Provider to assembly-ai.
Parameters reference
Universal-3.6 Pro parameters
Set these inside the assistant’stranscriber object. Vapi uses camelCase names; where a name differs from AssemblyAI’s streaming API parameter, the AssemblyAI name is given.
string
required
Set to
"assembly-ai".string
default:"universal-streaming-english"
The streaming model:
"universal-3-6-pro" (recommended), "universal-3-5-pro",
"universal-streaming-english", or "universal-streaming-multilingual". Vapi
defaults to "universal-streaming-english", so set it explicitly. See
Speech models.string
default:"balanced"
Accuracy/latency preset:
"min_latency", "balanced", or "max_accuracy".
Sets defaults for the mode-dependent fields on AssemblyAI’s side. Any value you
set explicitly, including minEndOfTurnSilenceWhenConfident and
maxTurnSilence, still takes precedence. Universal-3.6 Pro and 3.5 Pro only.
See Optimizing accuracy and
latency.number
AssemblyAI’s
min_turn_silence: milliseconds of silence before the model runs
an end-of-turn check. If the turn reads as complete, it ends. Otherwise the
turn stays open. When unset, Vapi doesn’t send it and the mode preset’s value applies (128 for
balanced). Only used when startSpeakingPlan.smartEndpointingPlan is not set.number
AssemblyAI’s
max_turn_silence: maximum milliseconds of silence before the
turn is forced to end, regardless of content. When unset, Vapi doesn’t send it
and the mode preset’s value applies (1280 for balanced). Turns that end
on a dictated entity (a phone number or email address) are held until this
limit, so it sets the response time for those turns. Only used when
startSpeakingPlan.smartEndpointingPlan is not set.boolean
default:"true"
A Vapi-side setting. When
true, Vapi holds AssemblyAI’s end-of-turn until its
own VAD also detects that the caller has stopped. When false, Vapi acts on
AssemblyAI’s end-of-turn immediately. We recommend false for faster
replies. Only used when startSpeakingPlan.smartEndpointingPlan is not set.string
Contextual prompt — a natural-language description of what the audio is about
(domain, scenario, or full details). Up to 1750 characters. Universal-3.6 Pro
and 3.5 Pro only. See Prompting.
string[]
Names, brands, or domain terms to boost. Up to 100 terms, each up to 50
characters. Included at no extra cost on Universal-3.6 Pro and 3.5 Pro. See
Key terms.
string
AssemblyAI’s
agent_context: context about what the assistant just said, used
to transcribe the next caller turn more accurately. Up to 1750 characters. Set
it to your firstMessage to seed the caller’s first reply. Universal-3.6 Pro
and 3.5 Pro only. See Conversation context.boolean
default:"false"
A Vapi-side setting. When
true, Vapi sends the text the assistant just spoke
to AssemblyAI as agent_context after every assistant turn. Off by default;
we recommend turning it on. Universal-3.6 Pro and 3.5 Pro only.string[]
AssemblyAI’s
language_codes: languages to steer transcription toward.
Universal-3.6 Pro code-switches across all 32 supported
languages by
default. Pass one code (["es"]) to pin a monolingual session, or several
(["en", "es"]) to constrain code-switching. Omit for automatic detection.
Universal-3.6 Pro and 3.5 Pro only.object
Transcribers to fall back to if the primary transcriber fails. See Fallback
transcribers.
AssemblyAITranscriber schema in Vapi’s Create assistant reference.
Some Universal-3.6 Pro parameters available on AssemblyAI’s API and in the LiveKit and Pipecat plugins aren’t exposed by Vapi’s transcriber, including
voice_focus, speaker_labels, interruption_delay, vad_threshold, and previous_context_n_turns. On Vapi, the model uses AssemblyAI’s defaults for these (or the values the mode preset sets).Vapi turn-taking parameters
These live on the assistant, outsidetranscriber, and shape how Vapi acts on AssemblyAI’s transcripts:
object
Vapi’s own endpointing (
"vapi", "livekit", or "custom-endpointing-model").
Leave unset to use AssemblyAI’s end-of-turn detection. When set, it
overrides AssemblyAI’s end-of-turn signal and the AssemblyAI silence fields
above are ignored. See Other turn detection
options.number
default:"0.4"
How long the assistant waits before speaking after the caller finishes, in
seconds (
0–5). With AssemblyAI’s end-of-turn detection, tune the AssemblyAI
silence thresholds and vadAssistedEndpointingEnabled first — they have a
larger effect on response time.object
When the caller’s speech interrupts the assistant:
numWords, voiceSeconds,
backoffSeconds, acknowledgementPhrases, and interruptionPhrases. See
Interruption handling.Legacy parameters
These apply to theuniversal-streaming-english and universal-streaming-multilingual models, but do not affect Universal-3.6 Pro or 3.5 Pro:
number
default:"0.7"
Confidence threshold for end-of-turn detection on the Universal-Streaming
models. Universal-3.6 Pro doesn’t use a confidence threshold to end turns, so
this field has no effect. Vapi’s AssemblyAI endpointing
presets
include it because they were written for Universal-Streaming.
boolean
default:"true"
Whether to return formatted final transcripts on the Universal-Streaming
models. Universal-3.6 Pro always returns formatted transcripts.
string
"en" or "multi", for the Universal-Streaming models. For Universal-3.6 Pro,
leave it unset and use languageCodes to steer languages.Turn detection
Vapi decides when the caller has finished speaking in one of two ways, chosen by whether the assistant has astartSpeakingPlan.smartEndpointingPlan:
- No
smartEndpointingPlan(recommended): Vapi uses AssemblyAI’s end-of-turn detection. - A
smartEndpointingPlan: Vapi’s own endpointing decides, and AssemblyAI provides the transcripts. See Other turn detection options.
AssemblyAI end-of-turn detection (recommended)
When to use: Most assistants. LeavesmartEndpointingPlan unset and Vapi acts on the end_of_turn flag AssemblyAI sends with each final transcript: it streams the turn to the LLM and generates a response.
Recommended starting configuration:
How it works:
- The caller speaks → Vapi streams the audio to AssemblyAI.
- The caller pauses for
minEndOfTurnSilenceWhenConfident(e.g.,128ms) → the model runs an end-of-turn check on what was said. - The turn doesn’t read as complete (no terminal punctuation) → a partial is emitted and the turn stays open.
- The turn ends in terminal punctuation, but the tail looks like an entity mid-dictation (an email address or phone number) → the turn is held anyway.
- The turn reads as complete → AssemblyAI ends the turn with
end_of_turn: true. - Silence reaches
maxTurnSilence(e.g.,1280ms) → the turn is forced closed regardless. - If
vadAssistedEndpointingEnabledistrue(Vapi’s default), Vapi also waits for its own VAD to register the caller as stopped before responding. Withfalse, Vapi responds as soon as AssemblyAI ends the turn.
mode preset. Values you set override the preset. Vapi’s transcriptionEndpointingPlan (onPunctuationSeconds, onNoPunctuationSeconds, onNumberSeconds) is not used in this mode; it only applies to transcribers without built-in end-of-turn detection.
Choosing a preset
Tuning maxTurnSilence
maxTurnSilence sets how long the assistant waits whenever a turn doesn’t read as complete:
- Dictated entities end at
maxTurnSilence. AssemblyAI holds a turn open when it ends on an entity such as a phone number or email address, even after terminal punctuation, so those turns close only when silence reachesmaxTurnSilence. Raising it protects entities, but it adds the same amount of wait to every entity turn. - Short, unpunctuated replies end at
maxTurnSilence. A bare “yes” or “okay” often comes back without terminal punctuation, so it also waits for the forced end. - Long pauses end the turn early. A caller who pauses longer than
maxTurnSilencemid-thought gets a reply before they finish. If your callers pause to look things up, raise it (e.g.,1500–2000ms).
Other turn detection options
SettingstartSpeakingPlan.smartEndpointingPlan hands turn-taking to Vapi. AssemblyAI keeps transcribing with the mode preset you set, but its end-of-turn signal no longer ends the turn, and minEndOfTurnSilenceWhenConfident, maxTurnSilence, and vadAssistedEndpointingEnabled are ignored. Use one of these options when you want Vapi to own turn-taking, for example to keep one endpointing configuration across several transcribers. See Vapi’s speech configuration guide for details.
Vapi smart endpointing
Vapi’s endpointing models evaluate the transcript to decide when the caller is done. Vapi recommends"livekit" for English and "vapi" for other languages:
"livekit", waitFunction maps the model’s probability that the caller is still speaking (x, from 0 to 1) to how many milliseconds to wait before replying. The value above is Vapi’s default. Lower coefficients reply sooner; higher ones wait longer when the model is unsure. When you evaluate either provider, include short answers such as “yes” in your tests.
Custom endpointing model
To decide the timeout yourself, set the provider to"custom-endpointing-model" and point it at your server:
call.endpointing.request message with the conversation so far each time a new transcript arrives, and your server responds with how long to wait:
server is omitted, Vapi sends the request to assistant.server, then org.server.
Custom endpointing rules
startSpeakingPlan.customEndpointingRules set a fixed timeout when a regex matches the caller’s speech, the assistant’s last message, or both. Vapi evaluates rules in order, and a matching rule takes precedence over the smartEndpointingPlan. For example, to wait longer after the assistant asks for a phone number:
type is "assistant" (match the assistant’s last message), "customer" (match the caller’s current transcript), or "both". timeoutSeconds ranges from 0 to 15.
Entity splitting tradeoff
Lower silence values produce faster turns but can split entities or utterances across turns. The two thresholds affect this differently.minEndOfTurnSilenceWhenConfident too low
The end-of-turn check fires during short pauses in a dictated entity:
maxTurnSilence too low
The forced end cuts the caller off mid-thought:
Universal-3.6 Pro’s formatting is significantly better when it has the
full entity in a single turn — email addresses, phone numbers, credit card
numbers, and physical addresses all benefit. If your assistant collects
alphanumeric input, raise
maxTurnSilence (e.g., to 2000–3000 ms) for that
assistant, or split entity collection into a dedicated assistant with a longer
window.Latency
A voice agent feels responsive when the gap between the caller finishing and the assistant replying is short. On Vapi, that gap is AssemblyAI’s end-of-turn timing plus Vapi’s own start-speaking settings. Start with themode preset — the highest-level dial for the accuracy/latency trade-off:
"min_latency" for the fastest turns or "max_accuracy" for the best transcripts. See Optimizing accuracy and latency.
From there, fine-tune the individual levers:
- End-of-turn timing.
minEndOfTurnSilenceWhenConfident(check) andmaxTurnSilence(forced end) directly control how soon AssemblyAI ends a turn. Lower is faster but risks splitting entities — see Entity splitting tradeoff. - VAD-assisted endpointing. Set
vadAssistedEndpointingEnabled: false. With Vapi’s default (true), Vapi waits for its own VAD to agree before acting on AssemblyAI’s end-of-turn, which adds time to every turn. - Vapi’s wait before speaking.
startSpeakingPlan.waitSeconds(default0.4s) is how long the assistant waits before speaking. With AssemblyAI’s end-of-turn detection, it has less effect on response time than the levers above. - Skip audio preprocessing. Noise suppression applied before audio reaches the model usually hurts accuracy more than the original noise. If you enable Vapi’s
backgroundSpeechDenoisingPlan, compare transcripts with and without it.
Latency breakdown
Accuracy
Universal-3.6 Pro is accurate out of the box. When you need more — domain vocabulary, proper nouns, known languages — reach for these levers. For entity-heavy dictation, also tune turn detection (see Entity splitting tradeoff).Prompting
Universal-3.6 Pro supports aprompt field for contextual prompting — a description of what the audio is about. Transcription behavior (verbatim output, punctuation, turn detection) is built in and optimized automatically; the prompt carries context, not instructions.
- Start with no prompt. Universal-3.6 Pro delivers strong accuracy without one — add context only if domain vocabulary is being misrecognized.
- Describe the conversation. Domain, scenario, or full details — start broad, and add only details your application actually knows (see the three context levels).
Key terms
UsekeytermsPrompt to boost recognition of specific names, brands, or domain terms — up to 100 terms, each up to 50 characters:
keytermsPrompt and prompt can be used together, with prompt describing the conversation and keytermsPrompt listing the terms.
Language selection
By default, Universal-3.6 Pro auto-detects and code-switches across all 32 supported languages. If you know which languages callers will speak,languageCodes steers transcription toward them for better accuracy on short or ambiguous turns:
Conversation context
Give the model both sides of the dialog so it transcribes the next caller turn more accurately. Universal-3.6 Pro keeps a short, per-session memory of the conversation:- The agent half — what the assistant just said, sent as
agentContext. - The caller half — prior finalized caller turns, carried forward automatically.
"What's your email address?", the model can produce "user@assemblyai.com" instead of "user at assemblyai dot com". This has the biggest impact on short replies ("yes", "7pm", single names) and spelled-out entities. See Conversation context for the full reference.
On Vapi, turn on agentContextAutoUpdateEnabled and seed the first turn with agentContext:
agentContextAutoUpdateEnabled is off by default on Vapi, unlike the
LiveKit and Pipecat integrations, where the agent’s replies are forwarded
automatically. When the caller interrupts a reply and Vapi detects the
interruption by voice activity (the default, stopSpeakingPlan.numWords: 0),
that interrupted reply isn’t sent.Per-call configuration
To tailor the transcriber to a single call, pass it inassistantOverrides on POST /call. This is useful when you know details before the call starts — the caller’s name, account number, or language:
provider and the fields that change, and the rest of the configuration (model, mode, silence thresholds) carries over. This is the opposite of PATCH, which replaces the whole object.
Interruption handling
Barge-in — the caller interrupting while the assistant is speaking — is handled by Vapi’sstopSpeakingPlan:
Choosing a strategy:
- Voice activity (
numWords: 0, the default). The fastest barge-in, driven by Vapi’s VAD. Background noise and short backchannels (“mhm”, “yeah”) can stop the assistant. RaisevoiceSeconds(up to0.5) if that happens too often. - Word count (
numWords: 1–3). The assistant stops only once AssemblyAI has transcribed that many words, and Vapi’sacknowledgementPhrasesfilter out backchannels. This resists false interruptions but depends on how quickly partial transcripts arrive. Themin_latencymode emits first partials sooner thanbalancedandmax_accuracy, so pair it with word-count interruption if barge-in feels slow — explicitminEndOfTurnSilenceWhenConfidentandmaxTurnSilencevalues still override the preset’s turn timing.
acknowledgementPhrases for your domain. In a booking flow, a bare “yes” is often a real confirmation, not a backchannel.
Fallback transcribers
If AssemblyAI is unavailable or rejects the session, Vapi can switch to another transcriber for the rest of the call. List fallbacks intranscriber.fallbackPlan.transcribers, in order:
fallbackPlan. When you create an assistant without a fallbackPlan, Vapi adds "autoFallback": {"enabled": true}, which lets it choose a fallback transcriber on its own; set an explicit fallbackPlan if you want to control which transcriber takes over. See Vapi’s transcriber fallback guide.
Speech models
Universal-3.6 Pro is recommended for all new Vapi assistants. The
Universal-Streaming models are maintained for backward compatibility but lack
prompting, conversation context, language steering, and the
mode preset.Billing
- Vapi-managed access (default). Vapi bills AssemblyAI transcription on your Vapi account. You don’t need an AssemblyAI account.
- Your own AssemblyAI key. Once you connect your key, AssemblyAI bills transcription on your AssemblyAI account and Vapi doesn’t charge for it. See AssemblyAI pricing.
Troubleshooting
Migrating from another transcriber
To balance accuracy, latency, turn-taking, and interruption handling, map your current Vapi configuration to AssemblyAI using the tables below.How is the assistant detecting end-of-turn today?
Which settings are you migrating from?
Migrating a production deployment? Talk to our team.