New Universal-3.6 Pro Realtime is now available Learn more
Insights & Use Cases

Multi-language voice agents: Building agents that speak to anyone

Multilingual voice agent guide: build real-time AI calls with STT, LLM, TTS, plus orchestration for fast language detection, switching, and accuracy at scale.

Abstract green half-sphere illustration

Written by

Kelsey Foster

Published on

30 September 2026

Most voice agents that claim to be multilingual are really monolingual agents with a language picker in front.

The caller chooses Spanish, the system loads a Spanish configuration, and everything downstream assumes Spanish for the rest of the call. It works right up until someone says their email address in English, or a bilingual speaker does what bilingual speakers do and switches halfway through a sentence.

Real multilingual behavior fails in a specific place: the transcription layer, on short conversational turns. This post covers what a multilingual voice agent is, where the languages live in the stack, and how to tell a good one from a demo. The code lives in the companion build tutorial.

What is a multilingual voice agent?

A multilingual voice agent is a conversational system that understands and responds in more than one language without being told in advance which one to expect, and without breaking when a speaker mixes languages inside a single turn. The second half is the part most systems can't do.

Three properties define one:

  • It doesn't need a language declared up front. No IVR menu, no routing by phone prefix, no separate agent per market.
  • It survives code-switching. "Mi número de cuenta es four-seven-two-alpha" comes back as one coherent turn, not two broken ones.
  • Its quality doesn't collapse outside English. Turn-taking, entity accuracy, and interruption handling hold up in every language it claims.

That last one is where most evaluations stop too early. Teams benchmark transcription accuracy per language, find the numbers acceptable, and ship — then discover the agent interrupts Japanese speakers constantly because its turn detection was tuned on English pause patterns. Conversational quality is language-dependent in ways that word error rate never shows you.

Which layer of the stack speaks which languages?

Language coverage isn't one number, because a voice agent isn't one component — and it isn't even one number per component, because understanding a language and speaking it are separate capabilities. The Voice Agent API understands 18 languages on input and speaks 6 of them back today. Your agent's real reach is the intersection of what each layer can do, in each direction.

Layer What it does Language coverage
Speech-to-text (Universal-3.6 Pro Realtime) Turns caller audio into text, detects turn ends, handles code-switching 32 languages, mid-sentence code-switching, Hinglish included
Voice Agent API (full pipeline) STT, LLM, and speech synthesis over a single WebSocket 18 input languages understood with native code-switching; 6 output languages spoken with native-accent voices
Bring-your-own stack Our speech model plus your own LLM and voice provider 32 on the speech side, whatever your other components support

The 18 input languages: English, Spanish, German, French, Portuguese, Italian, Turkish, Dutch, Swedish, Norwegian, Danish, Finnish, Hindi, Vietnamese, Arabic, Hebrew, Japanese, and Chinese. The streaming speech model, Universal-3.6 Pro Realtime, transcribes those 18 plus 14 more — Afrikaans, Cantonese, Catalan, Estonian, Galician, Korean, Marathi, Norwegian Nynorsk, Persian, Romanian, Russian, Urdu, Xhosa, and Zulu — for 32 in total. The 6 the agent can currently answer in with a native- or primary-accent voice: English, Italian, Spanish, German, Portuguese, and French. Native-accent voices for the remaining twelve are on the roadmap.

That asymmetry is worth being precise about, because it changes what you should build. If your callers speak inside the 18 and you need to answer inside the 6, the managed pipeline is one integration and one bill. If you need to answer in Hindi or Japanese today, you compose — our speech model underneath an orchestrator like LiveKit or Pipecat, both of which we ship drop-in plugins for, or alongside a platform like Vapi.

Notice which side the constraint sits on. Recognition is the solved half — the agent already understands nearly three times as many languages as it can speak. The boundary you're designing around is speech synthesis, not speech recognition, and it moves every time a new voice ships. Check your output coverage before you promise a market.

See how your languages perform

Coverage tables don't tell you how an agent sounds in Hindi or Portuguese. Get an API key and hear it on your own audio.

Sign up free

Why conversational speech breaks multilingual transcription

Voice agent audio is nothing like the audio speech models are usually benchmarked on, and multilingual audio makes the mismatch worse. Agent conversations are short, fragmentary, and full of exactly the tokens that carry all the meaning: a date, a name, a confirmation, six digits of an account number.

Think about what the model is being asked to do. Someone says "seven." One syllable, no surrounding words, possibly accented, possibly over a phone codec. Is that a quantity, an hour, a digit in a sequence, or a Spanish speaker saying something else entirely? There's no context in the audio because there's barely any audio.

But there is context in the conversation — the agent just asked a question. That's what agent_context is for. You pass the agent's spoken reply into the transcription session and the model knows what kind of answer to expect. Ask "what's your email address?" and a reply that would come back as "user at assemblyai dot com" resolves to "user@assemblyai.com". Short, fragmentary turns are exactly where that resolution pays off.

How much conversational context is worth depends on how much of it you can supply, and the streaming docs put published numbers on that. Benchmarked on 20,000 real voice-agent calls, these are the relative reductions against sending no prompt at all, at three levels of detail — domain only ("Medical consultation call."), a scenario sentence ("Cardiology consultation about chest pain symptoms."), and a full description naming people, products, and identifiers.

Improvement vs. no prompt Domain Scenario Detailed
Word error rate −5% −10% −21%
Hallucinated words −9% −12% −19%
Entity error rate — overall −2% −7% −29%
Entity error rate — names −5% −16% −49%
Entity error rate — places −9% −21% −44%
Entity error rate — medical terms −2% −24% −43%

Two things to take from that. Scenario context is the practical default — it asks for nothing your application doesn't already know, and it already cuts WER around 10% and name and place errors 16–21%. And detailed context is the upper bound worth reaching for when you have the information: name errors nearly halve. Counterintuitively, hallucinated words go down as you add context rather than up, and turn detection is unaffected. Full figures and the prompt-level definitions are in the streaming prompting documentation.

The other half comes free. Context Carryover keeps a short rolling memory of prior finalized turns, on by default and at no extra cost, so the user's side carries forward without you managing anything. The model launch post has the details.

One honest caveat: carrying prior turns biases the model toward languages already seen, and in sessions mixing three or more languages that can occasionally push it toward translating rather than transcribing. If you see that drift, pin a single transcription language in the prompt.

LiveKit, who make one of the orchestration frameworks a lot of these agents are built on, singled this out:

"We're excited to make AssemblyAI's Universal-3.5 Pro available on LiveKit Inference. What really stands out is their pace of innovation with Context Carryover — it intelligently applies conversation context to improve transcription accuracy in a way most speech models don't, removing the need for users to predefine key terms." — David Zhao, Co-founder at LiveKit

What does good look like on real agent conversations?

On AssemblyAI's English voice-agent benchmark, which scores models on 12,460 scripted voice-agent scenarios rather than read speech, Universal-3.6 Pro Realtime posts a 5.19% word error rate. Lower is better throughout, and the full methodology is in the Universal-3.6 Pro Realtime research write-up. Independent leaderboards point the same way: a 0.96% pooled semantic word error rate on Pipecat's open STT benchmark and the lowest word error rate on Coval's streaming leaderboard (2.2%).

Metric Universal-3.6 Pro Realtime Deepgram Flux EN ElevenLabs Scribe v2 Deepgram Nova-3
Word error rate 5.19% 13.50% 7.78% 8.64%
Entity error rate 14.4% 30.1% 18.5% 26.1%
Names 10.9% 29.0% 14.8% 24.3%
Codes / IDs 10.0% 46.1% 12.0% 27.6%
Phone numbers 2.4% 11.5% 3.4% 4.5%

Read the entity rows, not the WER row. Word error rate treats "the" and "Nakamura" as equally important, which is precisely backwards for an agent that has to look up a customer record. An entity error rate of 30.1% means nearly a third of the names, codes, and numbers come out wrong — and every one of those becomes a failed lookup, a wrong transfer, or a caller repeating themselves for the third time.

Those failures compound in a multilingual deployment, because entity errors cluster in exactly the places where languages meet: a Spanish speaker's English street name, a French caller spelling a German surname.

Designing the turn: pacing, noise, and the response budget

The hardest engineering problem in a multilingual voice agent isn't transcribing words. It's knowing when the caller has finished speaking.

Silence-based endpointing — wait 700ms, assume they're done — is a monolingual assumption dressed up as a technical setting. Speakers pause differently across languages, across dialects, and across how comfortable they are in the language they're using. A non-native speaker reaching for a word gets cut off by a threshold tuned on fluent English.

Universal-3.6 Pro Realtime decides a turn has ended by weighing whether what has been said so far reads as a completed thought, rather than endpointing on silence alone. You can also raise min_turn_silence mid-stream when you expect entity-like speech — someone spelling a name or reading digits — then re-send your mode preset afterward to restore fast endpointing. That matters for multilingual audio, where "they've stopped" and "they're thinking" sound very different depending on the language.

Noise is the other environmental variable, and it correlates with multilingual deployments more than people expect. voice_focus isolates the primary speaker and suppresses everything behind them, with near-field for headsets and phones and far-field for rooms and open spaces. Background speech in a second language is a particularly nasty failure mode, and this is what handles it.

End to end, the Voice Agent API runs around 1–1.3 seconds from the caller finishing to the agent starting, at a flat $4.50/hr covering speech-to-text, the language model, and speech synthesis together. One WebSocket, one bill, one set of logs, instead of three vendors to keep in sync. See full pricing.

Talk to an agent in your own language

Try code-switching, interruptions, and accented speech in the playground and see how the turn detection behaves.

Try playground

Where multilingual voice agents earn their keep

Four patterns come up repeatedly, and they have genuinely different accuracy requirements.

Contact center deflection in mixed-language markets. Miami, Toronto, Barcelona, Mumbai — places where the language of a call isn't predictable from the number dialled. The win isn't cost per call, it's removing the language-routing IVR that made callers pick a language before anyone had heard them speak. Contact center deployments are the most common starting point.

Front-line qualification and intake. Appointment booking, order status, delivery windows, insurance intake. Short calls dense with entities, which is why the entity rows in that benchmark table matter more here than the WER row.

Field and in-person voice. Drive-thrus, kiosks, warehouse floors. Far-field audio, high noise, and often a workforce and customer base that don't share a first language.

Language-learning agents. The one case where code-switching isn't an edge case but the entire product — a learner drops into their native language mid-sentence and the agent has to follow without losing the thread.

More in our guide to AI voice agents and across our voice agent solutions.

How to evaluate a multilingual voice agent before you ship

Aggregate word error rate is the least informative number you can collect about a multilingual agent. It averages across languages, across utterance lengths, and across token types, which hides every failure mode that actually matters. Test these five things instead.

  • Build code-switched test sets, not per-language ones. Clean Spanish calls and clean English calls tell you nothing about the calls containing both. Sample from real traffic and keep the mixing in.
  • Score entity accuracy per language. Names, dates, addresses, account numbers — broken out by language, not pooled. Pooled numbers look fine while your German name recognition quietly fails.
  • Measure turn-taking in each language separately. Count false interruptions and how long the agent waits after the caller genuinely stops. Both drift language to language, and both are far more visible to a caller than a misspelled word.
  • Test short utterances on purpose. One-and-two-word replies — confirmations, digits, single names — in every language you support. This is where agents fail and where aggregate metrics are blindest.
  • Test recovery, not just accuracy. When the agent mishears, does it ask a sensible clarifying question in the right language, or confidently proceed with the wrong value? Slightly worse accuracy with good recovery beats the reverse.

Our guide to choosing a speech-to-text API for voice agents goes deeper on methodology, and where voice agent stacks start showing their limits covers what breaks at scale.

Where the real gap is

Here's the thing that surprised us most in the voice agent benchmark data: the spread between the best and worst models on word error rate is about 8 points. On entity error rate it's nearly 16.

Which tells you the models aren't just differently accurate — they're differently accurate in a shape that matters. Getting "um" wrong costs nothing. Getting a surname wrong ends the call. Multilingual deployment amplifies that, because every language boundary is a place where entities get mangled, and every mangled entity is a caller who now wants a human.

The same lesson applies to the coverage question you started with. Eighteen languages in, six out, is not a recognition problem waiting to be solved — recognition is already three times ahead. It's a synthesis roadmap, and it tells you exactly which markets are one voice release away.

So evaluate for the shape, not the average. Then go build the thing — the step-by-step build guide linked above takes you from an empty file to a working multilingual agent.

Build your first multilingual agent

One WebSocket, one flat rate, no orchestration to maintain. Sign up and start a session in minutes.

Sign up free

Frequently asked questions

What is the best voice agent API for multilingual customer support agents?

Judge it on entity accuracy in your languages, not on aggregate word error rate. On AssemblyAI's English voice-agent benchmark, Universal-3.6 Pro Realtime posts a 5.19% WER and 14.4% entity error rate, compared with 8.64% and 26.1% for Deepgram Nova-3 and 13.50% and 30.1% for Deepgram Flux EN. For support agents specifically, the name and phone-number rows predict how often a caller has to repeat themselves.

Which voice agent API has the lowest latency for real-time conversation?

Latency numbers are only comparable if the turn detection is comparable, which is where most of the perceived delay comes from. The Voice Agent API runs around 1–1.3 seconds end to end, with end-of-turn detection that weighs whether what has been said reads as a completed thought rather than waiting out a silence timer. An agent with a slightly higher raw latency and better endpointing usually feels faster to a caller.

What languages does AssemblyAI's Voice AI support?

It depends on the layer and on the direction. Universal-3.6 Pro Realtime transcribes 32 languages with mid-sentence code-switching, including Hinglish, and the Voice Agent API understands 18 of them on input. On output it speaks 6 of them today with native-accent voices — English, Italian, Spanish, German, Portuguese, and French — with the remaining twelve on the roadmap. If you need to answer in a language outside that six right now, run our speech model underneath your own orchestration and voice provider.

Deepgram vs OpenAI Realtime vs AssemblyAI for voice agent speech-to-text — how do they differ?

They represent three different architectural bets. OpenAI Realtime uses one multimodal model for the whole pipeline; we use dedicated best-in-class models per step, which gives better turn detection and works out roughly 4x cheaper. Against Deepgram the difference is accuracy — 5.19% versus 13.50% (Flux EN) WER on AssemblyAI's English voice-agent benchmark — with no large upfront commitment required.

Can AssemblyAI handle multilingual audio with code-switching in a live conversation?

Yes. Universal-3.6 Pro Realtime handles mid-sentence language changes across its 32 languages in a single streaming session, with no language-pair configuration. Set an explicit language when you know it and leave it open when speakers genuinely mix.

Do I need a separate voice agent for each language I support?

No, and building one per language is the pattern worth avoiding. Separate agents mean the caller has to declare a language before speaking, and any caller who mixes languages breaks the routing. A single agent with a code-switching speech model handles all supported languages in one configuration.