Skip to main content

Overview

This guide covers configuring AssemblyAI’s Universal-3.6 Pro speech-to-text model as the transcriber for a Vapi voice assistant, using the Vapi API. Vapi is a hosted platform for building voice agents. You define an assistant — its transcriber, LLM, voice, and turn-taking behavior — as a JSON object, then connect it to phone numbers, web calls, or SIP. Vapi runs the real-time pipeline for you, so there’s no agent server to deploy.
Universal-3.6 Pro is our flagship next-generation streaming model for voice agents — multilingual and promptable, with conversation context.Set transcriber.provider to "assembly-ai" and transcriber.speechModel to "universal-3-6-pro". Vapi defaults to universal-streaming-english, so set the model explicitly.
AssemblyAI provides the speech-to-text and the turn detection in your Vapi assistant: Once you have an assistant running, tune it for what matters most to your use case:

Turn detection

Decide when the caller is done speaking, and how Vapi uses AssemblyAI’s end-of-turn signal.

Latency

Shorten the gap between the caller finishing and the assistant replying.

Accuracy

Prompting, key terms, language steering, and conversation context.

Interruptions

Barge-in and backchannels with Vapi’s stop speaking plan.
For a standalone voice agent without Vapi, see the AssemblyAI Voice Agent API, which handles STT, LLM routing, and TTS in a single WebSocket connection.

Quickstart

Create an assistant with Universal-3.6 Pro through the API, then call it.
1

Get your API keys

You need a Vapi private API key from the Vapi dashboard. Vapi’s API authenticates with it as a bearer token:
An AssemblyAI API key is optional. By default, Vapi provides access to AssemblyAI through its own account and bills transcription on your Vapi invoice. To use your own AssemblyAI account instead, see the next step.
2

Connect your AssemblyAI key (optional)

To have AssemblyAI bill transcription directly, add your AssemblyAI API key to Vapi. In the dashboard, open Integrations and add AssemblyAI, or create the credential through the API:
Once the key is validated, AssemblyAI transcription runs on your AssemblyAI account and Vapi no longer charges for it. See Vapi’s provider keys guide.
You can obtain an AssemblyAI API key by signing up for a free account and navigating to the API Keys tab of the dashboard.
3

Create an assistant

Create an assistant with a Universal-3.6 Pro transcriber. The request leaves startSpeakingPlan.smartEndpointingPlan unset, which is what tells Vapi to use AssemblyAI’s end-of-turn detection:
The response includes the assistant’s id. Always set voice explicitly: if you leave it out, Vapi assigns a default voice rather than the one you expect. The example uses Vapi’s built-in Elliot voice; see Vapi’s Create assistant reference for every voice provider and field.minEndOfTurnSilenceWhenConfident and maxTurnSilence above are the values the balanced preset uses, so they behave the same as mode alone; keeping them in the request makes the timing explicit and gives you the two values to tune. vadAssistedEndpointingEnabled: false makes Vapi act on AssemblyAI’s end-of-turn as soon as it arrives. See Turn detection for the recommended settings.
This guide calls the REST API directly. Vapi’s server SDKs (@vapi-ai/server-sdk, vapi_server_sdk) wrap the same endpoints, but their generated types can lag behind the API. If your SDK version’s types don’t include universal-3-6-pro yet, upgrade the SDK or send the request with the REST API.
4

Call the assistant

The quickest test is a web call from the dashboard: open Assistants, select your assistant, and select Talk.To have the assistant call a phone, pass its id and a Vapi phone number to POST /call:
To answer inbound calls, assign the assistant to a phone number in the dashboard. To embed the assistant in a web page, use the Vapi Web SDK. After each call, the call log in the dashboard shows the AssemblyAI transcript.

Switching an existing assistant to Universal-3.6 Pro

To move an existing assistant to AssemblyAI, update its transcriber with PATCH /assistant/{id}:
PATCH replaces the whole transcriber object. Fields you leave out are dropped, not kept. A patch that sends only {"provider": "assembly-ai", "keytermsPrompt": [...]} removes speechModel, and the assistant falls back to Vapi’s default model, universal-streaming-english. Always send the complete transcriber configuration.
If the assistant has a startSpeakingPlan.smartEndpointingPlan, remove it too, so Vapi uses AssemblyAI’s end-of-turn detection instead of its own. See Turn detection. In the dashboard, the same settings are on the assistant’s Transcriber tab: set Provider to assembly-ai.

Parameters reference

Universal-3.6 Pro parameters

Set these inside the assistant’s transcriber object. Vapi uses camelCase names; where a name differs from AssemblyAI’s streaming API parameter, the AssemblyAI name is given.
string
required
Set to "assembly-ai".
string
default:"universal-streaming-english"
The streaming model: "universal-3-6-pro" (recommended), "universal-3-5-pro", "universal-streaming-english", or "universal-streaming-multilingual". Vapi defaults to "universal-streaming-english", so set it explicitly. See Speech models.
string
default:"balanced"
Accuracy/latency preset: "min_latency", "balanced", or "max_accuracy". Sets defaults for the mode-dependent fields on AssemblyAI’s side. Any value you set explicitly, including minEndOfTurnSilenceWhenConfident and maxTurnSilence, still takes precedence. Universal-3.6 Pro and 3.5 Pro only. See Optimizing accuracy and latency.
number
AssemblyAI’s min_turn_silence: milliseconds of silence before the model runs an end-of-turn check. If the turn reads as complete, it ends. Otherwise the turn stays open. When unset, Vapi doesn’t send it and the mode preset’s value applies (128 for balanced). Only used when startSpeakingPlan.smartEndpointingPlan is not set.
number
AssemblyAI’s max_turn_silence: maximum milliseconds of silence before the turn is forced to end, regardless of content. When unset, Vapi doesn’t send it and the mode preset’s value applies (1280 for balanced). Turns that end on a dictated entity (a phone number or email address) are held until this limit, so it sets the response time for those turns. Only used when startSpeakingPlan.smartEndpointingPlan is not set.
boolean
default:"true"
A Vapi-side setting. When true, Vapi holds AssemblyAI’s end-of-turn until its own VAD also detects that the caller has stopped. When false, Vapi acts on AssemblyAI’s end-of-turn immediately. We recommend false for faster replies. Only used when startSpeakingPlan.smartEndpointingPlan is not set.
string
Contextual prompt — a natural-language description of what the audio is about (domain, scenario, or full details). Up to 1750 characters. Universal-3.6 Pro and 3.5 Pro only. See Prompting.
string[]
Names, brands, or domain terms to boost. Up to 100 terms, each up to 50 characters. Included at no extra cost on Universal-3.6 Pro and 3.5 Pro. See Key terms.
string
AssemblyAI’s agent_context: context about what the assistant just said, used to transcribe the next caller turn more accurately. Up to 1750 characters. Set it to your firstMessage to seed the caller’s first reply. Universal-3.6 Pro and 3.5 Pro only. See Conversation context.
boolean
default:"false"
A Vapi-side setting. When true, Vapi sends the text the assistant just spoke to AssemblyAI as agent_context after every assistant turn. Off by default; we recommend turning it on. Universal-3.6 Pro and 3.5 Pro only.
string[]
AssemblyAI’s language_codes: languages to steer transcription toward. Universal-3.6 Pro code-switches across all 32 supported languages by default. Pass one code (["es"]) to pin a monolingual session, or several (["en", "es"]) to constrain code-switching. Omit for automatic detection. Universal-3.6 Pro and 3.5 Pro only.
object
Transcribers to fall back to if the primary transcriber fails. See Fallback transcribers.
For every transcriber field, see the AssemblyAITranscriber schema in Vapi’s Create assistant reference.
Some Universal-3.6 Pro parameters available on AssemblyAI’s API and in the LiveKit and Pipecat plugins aren’t exposed by Vapi’s transcriber, including voice_focus, speaker_labels, interruption_delay, vad_threshold, and previous_context_n_turns. On Vapi, the model uses AssemblyAI’s defaults for these (or the values the mode preset sets).

Vapi turn-taking parameters

These live on the assistant, outside transcriber, and shape how Vapi acts on AssemblyAI’s transcripts:
object
Vapi’s own endpointing ("vapi", "livekit", or "custom-endpointing-model"). Leave unset to use AssemblyAI’s end-of-turn detection. When set, it overrides AssemblyAI’s end-of-turn signal and the AssemblyAI silence fields above are ignored. See Other turn detection options.
number
default:"0.4"
How long the assistant waits before speaking after the caller finishes, in seconds (0–5). With AssemblyAI’s end-of-turn detection, tune the AssemblyAI silence thresholds and vadAssistedEndpointingEnabled first — they have a larger effect on response time.
object
When the caller’s speech interrupts the assistant: numWords, voiceSeconds, backoffSeconds, acknowledgementPhrases, and interruptionPhrases. See Interruption handling.

Legacy parameters

These apply to the universal-streaming-english and universal-streaming-multilingual models, but do not affect Universal-3.6 Pro or 3.5 Pro:
number
default:"0.7"
Confidence threshold for end-of-turn detection on the Universal-Streaming models. Universal-3.6 Pro doesn’t use a confidence threshold to end turns, so this field has no effect. Vapi’s AssemblyAI endpointing presets include it because they were written for Universal-Streaming.
boolean
default:"true"
Whether to return formatted final transcripts on the Universal-Streaming models. Universal-3.6 Pro always returns formatted transcripts.
string
"en" or "multi", for the Universal-Streaming models. For Universal-3.6 Pro, leave it unset and use languageCodes to steer languages.

Turn detection

Vapi decides when the caller has finished speaking in one of two ways, chosen by whether the assistant has a startSpeakingPlan.smartEndpointingPlan:
  • No smartEndpointingPlan (recommended): Vapi uses AssemblyAI’s end-of-turn detection.
  • A smartEndpointingPlan: Vapi’s own endpointing decides, and AssemblyAI provides the transcripts. See Other turn detection options.
When to use: Most assistants. Leave smartEndpointingPlan unset and Vapi acts on the end_of_turn flag AssemblyAI sends with each final transcript: it streams the turn to the LLM and generates a response. Recommended starting configuration:
How it works:
  1. The caller speaks → Vapi streams the audio to AssemblyAI.
  2. The caller pauses for minEndOfTurnSilenceWhenConfident (e.g., 128 ms) → the model runs an end-of-turn check on what was said.
  3. The turn doesn’t read as complete (no terminal punctuation) → a partial is emitted and the turn stays open.
  4. The turn ends in terminal punctuation, but the tail looks like an entity mid-dictation (an email address or phone number) → the turn is held anyway.
  5. The turn reads as complete → AssemblyAI ends the turn with end_of_turn: true.
  6. Silence reaches maxTurnSilence (e.g., 1280 ms) → the turn is forced closed regardless.
  7. If vadAssistedEndpointingEnabled is true (Vapi’s default), Vapi also waits for its own VAD to register the caller as stopped before responding. With false, Vapi responds as soon as AssemblyAI ends the turn.
If you leave the two silence thresholds unset, Vapi doesn’t send them and AssemblyAI applies the values of the mode preset. Values you set override the preset. Vapi’s transcriptionEndpointingPlan (onPunctuationSeconds, onNoPunctuationSeconds, onNumberSeconds) is not used in this mode; it only applies to transcribers without built-in end-of-turn detection.

Choosing a preset

Tuning maxTurnSilence

maxTurnSilence sets how long the assistant waits whenever a turn doesn’t read as complete:
  • Dictated entities end at maxTurnSilence. AssemblyAI holds a turn open when it ends on an entity such as a phone number or email address, even after terminal punctuation, so those turns close only when silence reaches maxTurnSilence. Raising it protects entities, but it adds the same amount of wait to every entity turn.
  • Short, unpunctuated replies end at maxTurnSilence. A bare “yes” or “okay” often comes back without terminal punctuation, so it also waits for the forced end.
  • Long pauses end the turn early. A caller who pauses longer than maxTurnSilence mid-thought gets a reply before they finish. If your callers pause to look things up, raise it (e.g., 1500–2000 ms).
See Entity splitting tradeoff for examples.
Check what Vapi sent to AssemblyAI. Open the call in the Vapi dashboard and find the AssemblyAI Universal WebSocket request entry in its logs. The request URL lists the exact speech_model, mode, min_turn_silence, and max_turn_silence for the call.

Other turn detection options

Setting startSpeakingPlan.smartEndpointingPlan hands turn-taking to Vapi. AssemblyAI keeps transcribing with the mode preset you set, but its end-of-turn signal no longer ends the turn, and minEndOfTurnSilenceWhenConfident, maxTurnSilence, and vadAssistedEndpointingEnabled are ignored. Use one of these options when you want Vapi to own turn-taking, for example to keep one endpointing configuration across several transcribers. See Vapi’s speech configuration guide for details.

Vapi smart endpointing

Vapi’s endpointing models evaluate the transcript to decide when the caller is done. Vapi recommends "livekit" for English and "vapi" for other languages:
With "livekit", waitFunction maps the model’s probability that the caller is still speaking (x, from 0 to 1) to how many milliseconds to wait before replying. The value above is Vapi’s default. Lower coefficients reply sooner; higher ones wait longer when the model is unsure. When you evaluate either provider, include short answers such as “yes” in your tests.

Custom endpointing model

To decide the timeout yourself, set the provider to "custom-endpointing-model" and point it at your server:
Vapi sends your server a call.endpointing.request message with the conversation so far each time a new transcript arrives, and your server responds with how long to wait:
If server is omitted, Vapi sends the request to assistant.server, then org.server.

Custom endpointing rules

startSpeakingPlan.customEndpointingRules set a fixed timeout when a regex matches the caller’s speech, the assistant’s last message, or both. Vapi evaluates rules in order, and a matching rule takes precedence over the smartEndpointingPlan. For example, to wait longer after the assistant asks for a phone number:
type is "assistant" (match the assistant’s last message), "customer" (match the caller’s current transcript), or "both". timeoutSeconds ranges from 0 to 15.

Entity splitting tradeoff

Lower silence values produce faster turns but can split entities or utterances across turns. The two thresholds affect this differently.

minEndOfTurnSilenceWhenConfident too low

The end-of-turn check fires during short pauses in a dictated entity:

maxTurnSilence too low

The forced end cuts the caller off mid-thought:
Universal-3.6 Pro’s formatting is significantly better when it has the full entity in a single turn — email addresses, phone numbers, credit card numbers, and physical addresses all benefit. If your assistant collects alphanumeric input, raise maxTurnSilence (e.g., to 2000–3000 ms) for that assistant, or split entity collection into a dedicated assistant with a longer window.

Latency

A voice agent feels responsive when the gap between the caller finishing and the assistant replying is short. On Vapi, that gap is AssemblyAI’s end-of-turn timing plus Vapi’s own start-speaking settings. Start with the mode preset — the highest-level dial for the accuracy/latency trade-off:
Use "min_latency" for the fastest turns or "max_accuracy" for the best transcripts. See Optimizing accuracy and latency. From there, fine-tune the individual levers:
  • End-of-turn timing. minEndOfTurnSilenceWhenConfident (check) and maxTurnSilence (forced end) directly control how soon AssemblyAI ends a turn. Lower is faster but risks splitting entities — see Entity splitting tradeoff.
  • VAD-assisted endpointing. Set vadAssistedEndpointingEnabled: false. With Vapi’s default (true), Vapi waits for its own VAD to agree before acting on AssemblyAI’s end-of-turn, which adds time to every turn.
  • Vapi’s wait before speaking. startSpeakingPlan.waitSeconds (default 0.4 s) is how long the assistant waits before speaking. With AssemblyAI’s end-of-turn detection, it has less effect on response time than the levers above.
  • Skip audio preprocessing. Noise suppression applied before audio reaches the model usually hurts accuracy more than the original noise. If you enable Vapi’s backgroundSpeechDenoisingPlan, compare transcripts with and without it.

Latency breakdown

Accuracy

Universal-3.6 Pro is accurate out of the box. When you need more — domain vocabulary, proper nouns, known languages — reach for these levers. For entity-heavy dictation, also tune turn detection (see Entity splitting tradeoff).

Prompting

Universal-3.6 Pro supports a prompt field for contextual prompting — a description of what the audio is about. Transcription behavior (verbatim output, punctuation, turn detection) is built in and optimized automatically; the prompt carries context, not instructions.
Tips:
  • Start with no prompt. Universal-3.6 Pro delivers strong accuracy without one — add context only if domain vocabulary is being misrecognized.
  • Describe the conversation. Domain, scenario, or full details — start broad, and add only details your application actually knows (see the three context levels).

Key terms

Use keytermsPrompt to boost recognition of specific names, brands, or domain terms — up to 100 terms, each up to 50 characters:
keytermsPrompt and prompt can be used together, with prompt describing the conversation and keytermsPrompt listing the terms.

Language selection

By default, Universal-3.6 Pro auto-detects and code-switches across all 32 supported languages. If you know which languages callers will speak, languageCodes steers transcription toward them for better accuracy on short or ambiguous turns:
See Multilingual transcription for the full language list.

Conversation context

Give the model both sides of the dialog so it transcribes the next caller turn more accurately. Universal-3.6 Pro keeps a short, per-session memory of the conversation:
  • The agent half — what the assistant just said, sent as agentContext.
  • The caller half — prior finalized caller turns, carried forward automatically.
With the assistant’s question in context, the model can anticipate the answer, sharpen entity recognition, and disambiguate similar-sounding words. For example, after the assistant asks "What's your email address?", the model can produce "user@assemblyai.com" instead of "user at assemblyai dot com". This has the biggest impact on short replies ("yes", "7pm", single names) and spelled-out entities. See Conversation context for the full reference. On Vapi, turn on agentContextAutoUpdateEnabled and seed the first turn with agentContext:
agentContextAutoUpdateEnabled is off by default on Vapi, unlike the LiveKit and Pipecat integrations, where the agent’s replies are forwarded automatically. When the caller interrupts a reply and Vapi detects the interruption by voice activity (the default, stopSpeakingPlan.numWords: 0), that interrupted reply isn’t sent.

Per-call configuration

To tailor the transcriber to a single call, pass it in assistantOverrides on POST /call. This is useful when you know details before the call starts — the caller’s name, account number, or language:
Overrides merge with the assistant’s saved transcriber: pass provider and the fields that change, and the rest of the configuration (model, mode, silence thresholds) carries over. This is the opposite of PATCH, which replaces the whole object.

Interruption handling

Barge-in — the caller interrupting while the assistant is speaking — is handled by Vapi’s stopSpeakingPlan:
Choosing a strategy:
  • Voice activity (numWords: 0, the default). The fastest barge-in, driven by Vapi’s VAD. Background noise and short backchannels (“mhm”, “yeah”) can stop the assistant. Raise voiceSeconds (up to 0.5) if that happens too often.
  • Word count (numWords: 1–3). The assistant stops only once AssemblyAI has transcribed that many words, and Vapi’s acknowledgementPhrases filter out backchannels. This resists false interruptions but depends on how quickly partial transcripts arrive. The min_latency mode emits first partials sooner than balanced and max_accuracy, so pair it with word-count interruption if barge-in feels slow — explicit minEndOfTurnSilenceWhenConfident and maxTurnSilence values still override the preset’s turn timing.
Edit acknowledgementPhrases for your domain. In a booking flow, a bare “yes” is often a real confirmation, not a backchannel.

Fallback transcribers

If AssemblyAI is unavailable or rejects the session, Vapi can switch to another transcriber for the rest of the call. List fallbacks in transcriber.fallbackPlan.transcribers, in order:
Fallback entries take the same fields as the primary transcriber, except fallbackPlan. When you create an assistant without a fallbackPlan, Vapi adds "autoFallback": {"enabled": true}, which lets it choose a fallback transcriber on its own; set an explicit fallbackPlan if you want to control which transcriber takes over. See Vapi’s transcriber fallback guide.

Speech models

Universal-3.6 Pro is recommended for all new Vapi assistants. The Universal-Streaming models are maintained for backward compatibility but lack prompting, conversation context, language steering, and the mode preset.

Billing

  • Vapi-managed access (default). Vapi bills AssemblyAI transcription on your Vapi account. You don’t need an AssemblyAI account.
  • Your own AssemblyAI key. Once you connect your key, AssemblyAI bills transcription on your AssemblyAI account and Vapi doesn’t charge for it. See AssemblyAI pricing.
Telephony, LLM, and TTS usage are billed separately either way.

Troubleshooting

Migrating from another transcriber

To balance accuracy, latency, turn-taking, and interruption handling, map your current Vapi configuration to AssemblyAI using the tables below.

How is the assistant detecting end-of-turn today?

Which settings are you migrating from?

Migrating a production deployment? Talk to our team.

Resources