Real-time transcription that code-switches for multilingual speakers
Learn how Universal-Streaming enables real-time transcription for multilingual speakers who naturally code-switch mid-sentence, handling six languages in a single forward pass. Built for hybrid-language conversations like bilingual customer calls and international meetings, it delivers low-latency, highly accurate transcripts without added complexity.



In Miami, business calls move between English and Spanish inside a single thought. In Montreal, meetings shift between French and English depending on who's speaking and what the term of art is. In Mumbai and every Indian diaspora city, Hinglish blends English vocabulary into Hindi grammar so thoroughly that the two aren't separable at the sentence level.
A sentence might start in one language, borrow a phrase from another, and switch back inside a few seconds. Say this out loud and it sounds completely normal:
"I am traveling to Arizona and I want to comprar unos tiquets de avión."
Nothing about that is unusual to the person saying it. To a conventional transcription pipeline, it's a disaster.
Per-turn language detection is not code-switching support
This is the distinction that matters most, and it's the one most vendors blur.
The common architecture puts a language identification model in front of the recognizer. Audio arrives, the gateway classifies it as Spanish, and routes it to the Spanish model. That works when someone speaks one language per turn. It fails the moment a language changes mid-utterance, because the routing decision was already made — the Spanish model gets English audio and transcribes it as the nearest Spanish-sounding nonsense.
It also costs you time. A detection gateway is a stage, and stages add latency to a path where you're counting milliseconds.
Universal-3.6 Pro Realtime has no language gateway. All 32 languages are handled in a single unified model, in one forward pass, so a switch inside a clause is just more audio — not a routing event. There's nothing to detect, nothing to re-route, and nothing to restart.
| Approach | Handles mid-sentence switching | Added latency | Failure mode |
|---|---|---|---|
| Language detection gateway | No | A full classification stage before recognition | Wrong-language transcription of the switched span |
| Per-turn language selection | Only between turns | Re-selection cost at each turn | Degrades on every intra-turn switch |
| Single unified multilingual model | Yes | None — one forward pass | Accuracy varies by language pair, not by switch point |
The 32 languages, and what "32" actually buys you
Universal-3.6 Pro Realtime streams these 32 languages with native code-switching between them:
English · Spanish · French · German · Italian · Portuguese · Arabic · Danish · Dutch · Finnish · Hebrew · Hindi · Japanese · Chinese · Norwegian · Swedish · Turkish · Vietnamese · Afrikaans · Cantonese · Catalan · Estonian · Galician · Korean · Marathi · Norwegian Nynorsk · Persian · Romanian · Russian · Urdu · Xhosa · Zulu
The count is worth reading carefully, because language counts are the most abused number in this industry. Thirty-two here does not mean 32 separate models you pick between. It means one model that hears all 32 at once, so any pair among them can alternate inside a sentence without you configuring anything. English–Spanish, French–English, Hindi–English, Arabic–French: all the same operation to the model.
Hinglish is the sharpest test case. English nouns and verbs sit inside Hindi sentence structure so densely that "which language is this" isn't a well-formed question about a given clause. A router has to pick one and be wrong about half the words. A unified model doesn't have to pick.
For pre-recorded audio, coverage extends further — anything outside Universal-3.5 Pro's 18 pre-recorded languages falls back automatically to Universal-2 for 99 languages total. That fallback is a pre-recorded behavior. On streaming, 32 is 32. We wrote more about the broader picture in our guide to the multilingual speech-to-text API.
Speak two languages in one sentence and watch the transcript keep up. Test streaming transcription speed and accuracy on your own audio.
Why entity accuracy is the number that matters here
Ask a vendor about multilingual accuracy and you'll get a word error rate. It's the wrong question.
WER treats every word as equally important, which means "um" and a customer's account number carry identical weight. In a real multilingual conversation, the words that switch languages are almost never filler — they're product names, place names, technical terms, numbers. Those are exactly the words a WER average smooths over, and exactly the words your application depends on. We made the longer argument in why word error rate is broken.
So look at entity accuracy. On AssemblyAI's English voice-agent benchmark, which uses 12,460 scripted voice-agent scenarios rather than clean read speech, Universal-3.6 Pro Realtime posts these numbers:
| Metric (lower is better) | Universal-3.6 Pro Realtime | Deepgram Flux EN | ElevenLabs Scribe v2 | Deepgram Nova-3 |
|---|---|---|---|---|
| Entity error rate | 14.4% | 30.1% | 18.5% | 26.1% |
| Names | 10.9% | 29.0% | 14.8% | 24.3% |
| Codes / IDs | 10.0% | 46.1% | 12.0% | 27.6% |
| Phone numbers | 2.4% | 11.5% | 3.4% | 4.5% |
| Word error rate | 5.19% | 13.50% | 7.78% | 8.64% |
Note the shape of that table. Deepgram Nova-3 sits within about three and a half points of us on WER, yet its entity error rate is nearly double ours. If you'd picked a provider on the WER column alone, you'd have made a materially worse decision. More on why in our piece on entity accuracy in speech-to-text.
On code-switching specifically, Universal-3.6 Pro Realtime posts a 7.20% word error rate across five English language pairs on the same benchmark, down from 8.65% for Universal-3.5 Pro and ahead of ElevenLabs Scribe v2 (13.01%) and Deepgram Nova-3 Multi (17.13%). Universal-3.5 Pro had already delivered a 22% relative reduction in word error rate on code-switched audio, with a further 4% when you supply prompts; the launch write-up on code-switching and contextual prompting has the detail.
The latency picture: on Pipecat's open STT benchmark, the final transcript lands a median 307 ms after the speaker stops talking, alongside a 0.96% pooled semantic word error rate. Multilingual audio doesn't cost you latency here — there's no detection stage to pay for.
How to stream multilingual audio
You connect to the streaming WebSocket at wss://streaming.assemblyai.com/v3/ws and pass your connection parameters:
RATE = 16000
BASE_URL = "wss://streaming.assemblyai.com/v3/ws"
CONNECTION_PARAMS = {
"sample_rate": RATE,
"speech_model": "universal-3-6-pro",
"language_codes": ["en", "es"],
}
Four things worth calling out.
speech_model is singular on streaming. Pre-recorded transcription uses speech_models, a plural ordered array. Streaming takes one model name. Mixing them up is the single most common integration error we see.
Set the model explicitly. A config that omits it rides whatever the account default happens to be — exactly the kind of thing that quietly changes underneath you.
language_codes is plural, and it's a bias, not a filter. Pass the languages you actually expect in the conversation. It steers the model toward those languages without locking out the others, which is what makes a switch into an unexpected third language degrade gracefully instead of catastrophically. Note the plural: singular language_code is the pre-recorded parameter, and Universal Streaming's older language parameter is deprecated. Pass it explicitly rather than relying on any default. If you've fought with language steering before, our post on fixing language steering in streaming transcription covers the failure patterns.
Formatting is always on. Punctuation, capitalization and intelligent endpointing come standard, so "I'm going to la tienda" arrives formatted rather than as a wall of lowercase text. Model selection details are in the streaming model selection docs.
Keyterms prompting, at no extra cost
Multilingual audio is where domain vocabulary breaks hardest, because a product name pronounced with a Spanish phonology and transcribed by an English-leaning decoder can land anywhere. Keyterms prompting fixes that: pass up to 100 terms, each 50 characters or under, and update them mid-stream with UpdateConfiguration as the conversation moves.
On the streaming flagship, keyterms are included at no additional cost. On both streaming and pre-recorded audio you can combine general prompting with keyterms. The prompting and keyterms documentation covers both.
Context is what makes code-switching work
Here's the part that surprised us internally: the biggest accuracy gains on multilingual audio didn't come from more multilingual training data. They came from telling the model what the conversation is about.
agent_context lets you pass the question your agent just asked, so the model hears the reply through that lens. If the agent asked "what's your account number," a stream of digits gets interpreted as digits rather than as similar-sounding words in whichever language the caller reaches for. Across a benchmark of 20,000 voice agent audio files, passing agent context cut word error rate by 10.2%.
Context Carryover is rolling conversation memory. It's on by default and there's nothing to configure — earlier turns inform later ones, so a name established in minute one is still recognized in minute nine, even if the speaker switches languages in between. This matters disproportionately for code-switching, because the same entity often appears with different phonology in each language.
David Zhao, Co-founder at LiveKit, put the value plainly:
We're excited to make AssemblyAI's Universal-3.5 Pro available on LiveKit Inference. What really stands out is their pace of innovation with Context Carryover — it intelligently applies conversation context to improve transcription accuracy in a way most speech models don't, removing the need for users to predefine key terms.
More on the mechanics in our write-ups on conversation context for voice agents and agent context and carryover on LiveKit.
Evaluate real-time speech-to-text with low latency and strong accuracy. Launch pilots quickly with clear docs and developer-friendly APIs.
What it costs
Universal-3.6 Pro Realtime is $0.45 per hour — $0.0075 per minute, about three-quarters of a cent. Every language costs the same. There is no premium for non-English transcription and no surcharge for code-switched audio.
One billing detail people miss: streaming is billed on session duration, meaning how long the WebSocket stays open, not on how much audio contains speech. Close idle sockets.
| Item | Price | Notes |
|---|---|---|
| Universal-3.6 Pro Realtime | $0.45/hr | All 32 languages, same rate |
| Keyterms prompting | Included | Up to 100 terms, updatable mid-stream |
| General prompting | +$0.05/hr | Combinable with keyterms on streaming |
| Diarization with revision | +$0.12/hr | Live labels, up to 10 speakers |
| Voice Focus | +$0.10/hr | Near-field or far-field noise handling |
The free tier includes 333 hours of streaming transcription and 185 hours of pre-recorded transcription, which is enough to run a real pilot rather than a demo. Everything is pay-as-you-go, billed per second, with unlimited concurrency and no upfront commitment. The full itemized rate card lives on our pricing page.
Where this shows up in production
Voice agents serving multilingual markets. A Spanish-speaking customer switches to English for a technical term and back again. With a routing architecture, that's where the agent loses the thread. Our guide to building a multilingual voice agent walks the full stack.
Real-time agent assist for multilingual support teams. Suggestions surfaced to a human agent are only useful if the transcript underneath them is right, and support conversations in bilingual markets code-switch constantly.
Meeting assistants across international offices. Teams that share a second language switch into it for speed and precision. A transcript that garbles those spans loses exactly the parts people cared enough to say precisely.
Clinical documentation. Patient consultations in multilingual communities routinely mix languages — often the clinician in one, the patient in another, an interpreter bridging both. Two caveats matter here. Medical Mode, which handles clinical terminology, covers English, Spanish, German and French across pre-recorded and streaming — four languages, not 32. And on the compliance side, the accurate wording is this:
AssemblyAI enables covered entities and their business associates subject to HIPAA to use the AssemblyAI services to process protected health information (PHI). AssemblyAI is considered a business associate under HIPAA, and we offer a standard Business Associate Addendum (BAA) that is required under HIPAA to ensure that AssemblyAI appropriately safeguards PHI.
The BAA can be signed in minutes without a sales call, and PHI redaction is available across both audio and transcripts. See our medical transcription solutions for the full picture.
Fireflies evaluated the model for their voice agent pipeline. Foysal Osmany, Software Engineer at Fireflies:
We were searching for the best realtime ASR model for our voice agent pipeline in Fireflies. The new Universal 3.5 Pro speech model from Assembly is best so far in terms of accuracy, latency and language switching.
Getting started
Set speech_model to universal-3-6-pro, pass the language_codes you expect, open the socket, and start streaming. That's the whole integration. You can also test it without writing anything: open the playground, pick streaming, and speak two languages in one sentence.
But here's the thing worth taking away, and it isn't about language counts.
The industry spent years treating multilingual support as a routing problem — detect the language, dispatch to the right model, stitch the results. That framing produced systems that were fine at handling many languages and bad at handling mixed ones, which is backwards, because mixed is what people actually speak. Collapsing detection and recognition into one pass didn't just remove a latency stage. It removed the assumption that a span of audio has exactly one language, and that assumption was the bug the whole time.
Build for the way your users talk, not the way language dropdowns imply they should.
Learn how automation reduces costs, shortens handle times, and scales support without adding headcount. Get guidance tailored to your industry and goals.
Frequently asked questions
Can AssemblyAI handle multilingual audio with code-switching?
Yes — Universal-3.6 Pro Realtime handles mid-sentence code-switching natively across 32 languages, including Hinglish, in a single forward pass with no language detection gateway. On code-switched audio it posts a 7.20% word error rate across five English language pairs on AssemblyAI's voice-agent benchmark, down from 8.65% for Universal-3.5 Pro. You enable it by setting speech_model to universal-3-6-pro and passing the languages you expect in language_codes. Because there's no routing stage, a switch inside a clause costs nothing in latency.
What languages does AssemblyAI's Voice AI support?
Streaming with Universal-3.6 Pro Realtime supports 32 languages: English, Spanish, French, German, Italian, Portuguese, Arabic, Danish, Dutch, Finnish, Hebrew, Hindi, Japanese, Chinese, Norwegian, Swedish, Turkish, Vietnamese, Afrikaans, Cantonese, Catalan, Estonian, Galician, Korean, Marathi, Norwegian Nynorsk, Persian, Romanian, Russian, Urdu, Xhosa and Zulu. For pre-recorded audio, coverage extends to 99 languages in total — Universal-3.5 Pro handles 18 languages, and anything outside them falls back automatically to Universal-2. All 32 streaming languages are priced identically at $0.45/hr, with no premium for non-English audio.
How do transcription services handle multiple languages in the same audio?
Most services put a language identification model in front of the recognizer and route each turn to a language-specific model, which works for one-language-per-turn speech and fails on mid-sentence switching, because the routing decision is already made when the language changes. A unified multilingual model takes a different approach: all supported languages are recognized in one pass, so a switch inside a clause is just more audio rather than a routing event. That's how Universal-3.6 Pro Realtime works, which is why it adds no detection stage to the latency path and delivers final transcripts a median 307 ms after the speaker stops on Pipecat's open benchmark.
Does AssemblyAI support real-time streaming transcription?
Yes. Streaming runs over a WebSocket at wss://streaming.assemblyai.com/v3/ws with Universal-3.6 Pro Realtime, at $0.45/hr billed per second on session duration. On Pipecat's open STT benchmark, the final transcript lands a median 307 ms after the speaker stops talking. The free tier includes 333 hours of streaming transcription, and concurrency is unlimited with no upfront commitment.
What streaming features does AssemblyAI support?
Streaming includes 32-language code-switching, keyterms prompting at no extra cost (up to 100 terms, updatable mid-stream via UpdateConfiguration), agent_context for passing the agent's question to the model, Context Carryover for rolling conversation memory, and always-on punctuation, capitalization and intelligent endpointing. Optional add-ons include speaker diarization with revision at +$0.12/hr for up to 10 speakers, Voice Focus at +$0.10/hr for near-field or far-field noise handling, and general prompting at +$0.05/hr. Turn detection is configurable through min_turn_silence and max_turn_silence, whose defaults come from the mode preset (128 ms and 1280 ms on balanced), and vad_threshold (0.2 default), with three latency modes available: min_latency, balanced and max_accuracy.
AssemblyAI vs Speechmatics for multilingual transcription: how should I compare them?
Compare on entity accuracy for your own audio rather than on published word error rates or language counts, because those two numbers hide the failures that break applications. On AssemblyAI's English voice-agent benchmark, Universal-3.6 Pro Realtime posts a 14.4% entity error rate, 10.9% on names, 10.0% on codes and IDs and 2.4% on phone numbers, with a 5.19% WER; on Pipecat's open STT benchmark, final transcripts land a median 307 ms after the speaker stops, across 32 streaming languages at $0.45/hr. When you evaluate any provider, run audio that actually code-switches mid-sentence — not one-language-per-turn samples — since that's where routing-based architectures fail and unified models don't. Our full methodology and results are on the benchmarks page.
Lorem ipsum dolor sit amet, consectetur adipiscing elit, sed do eiusmod tempor incididunt ut labore et dolore magna aliqua. Ut enim ad minim veniam, quis nostrud exercitation ullamco laboris nisi ut aliquip ex ea commodo consequat. Duis aute irure dolor in reprehenderit in voluptate velit esse cillum dolore eu fugiat nulla pariatur.




