Your agent can take no for an answer
Universal-3.6 Pro Realtime is live. The same accuracy 3.5 Pro had on quiet calls, now on loud rooms, accents and phone lines, where every agent gets tested.
Today we’re releasing Universal-3.6 Pro Realtime. It is the same model line you’re running
today, trained on tens of thousands of hours of real voice-agent and telephony conversations, so the
accuracy 3.5 Pro delivered on quiet calls now holds on loud rooms, accents and phone lines. It serves 32
languages from one model with automatic language detection, runs at the same latency as 3.5 Pro, and is
available now as universal-3-6-pro at
$0.45/hr.
Everything below was measured against Universal-3.5 Pro on our English voice-agent benchmark: 12,460 scripted caller scenarios voiced by real speakers, on matched hardware at the same latency operating point. Every number comes with its base rate.
| English voice-agent benchmark | Universal-3.5 Pro | Universal-3.6 Pro |
|---|---|---|
| Short response error rate | 2.65% | 1.45% |
| Normalized word error rate | 5.80% | 5.13% |
| Word error rate | 13.59% | 12.32% |
| Language confusion (share of sessions) | 0.55% | 0.14% |
| Background-speech leakage (words per 1,000) | 3.53 | 2.52 |
Word error rate is given twice: the normalized rate compares transcripts after normalizing case, punctuation, numbers and abbreviations, and the raw rate also counts those formatting differences, which is why it is higher. The research post has the full methodology.
Short answers, heard right in the rooms callers are actually in
The turn that decides a call is usually one word: yes, no thanks, that’s it. When the recognizer drops or flips it, the agent books the wrong appointment, and the failure shows up as a refund or an escalation rather than a transcript error anyone reviews. Most teams cover this with a confirmation loop, “just to confirm, did you say yes?”, which adds a turn to every decision.
Universal-3.6 Pro gets those turns wrong 1.45% of the time, down from 2.65% on 3.5 Pro. That is 45% fewer wrong confirmations overall, and 52% fewer in heavy background noise. The metric scores by meaning: “yeah” for “yes” is not an error; a different answer or a dropped answer is.
| Short response error rate | Universal-3.5 Pro | Universal-3.6 Pro |
|---|---|---|
| No background noise | 1.84% | 1.14% |
| Low background noise | 2.69% | 1.55% |
| Heavy background noise | 3.44% | 1.66% |
Half as many wrong confirmations is the difference between a confirmation loop on every decision and a confirmation loop on the ones that carry money.
The gain tracks how hard the audio is. On clean recordings the rate falls 38%, on light background noise 42%, and on the heavy-noise third 52%. Pauses and hesitations are silence the model has to leave alone, and it now does.
Background voices stay out of the transcript
Background speech leakage counts words transcribed from a television, a coworker or the next desk and handed to your agent as if the caller said them. It falls from 3.53 to 2.52 words per 1,000 overall, and from 6.82 to 4.69 on the heavy-noise third of the benchmark. On clean audio both models are already near zero, so the gain lands on the calls that need it.
| Background-speech leakage (words per 1,000) | Universal-3.5 Pro | Universal-3.6 Pro |
|---|---|---|
| No background noise | 0.98 | 0.81 |
| Low background noise | 2.82 | 2.08 |
| Heavy background noise | 6.82 | 4.69 |
Voice focus changed underneath. Voice focus is an optional stage that suppresses speakers farther from the microphone than the caller. In 3.5 Pro the enhanced audio only decided where turns begin and end; the model still transcribed the original signal, so a second talker inside the caller’s turn came through. In 3.6 Pro the enhanced audio reaches the recognition model. On the ai-coustics test-call set, clips cut from real calls with a competing talker, word error rate against the primary speaker falls from 53.8% to 20.2% with voice focus on at its default setting.
Most teams already run something in front of the recognizer: a noise suppressor, a speaker-isolation stage, usually tuned for what sounds clean to a person on the other end of a call. That is a different objective from what a recognizer needs, and a filter that scrubs audio for human ears can take the cues the model was using with it. Voice focus is not a stage in front of the model. The filter and the recognition model ship together and are measured together, so what the filter removes shows up as word error rate rather than as a cleaner-sounding recording.
Voice focus is for other people’s voices, and it is off by default. Where the challenge is background noise rather than other talkers, 3.6 Pro handles it on its own, so turn voice focus on only when a second person is talking near the microphone and let the model do the rest.
It stays in the caller’s language, across 32 of them
A caller switches to Spanish mid-sentence, or has an accent, and 3.5 Pro occasionally committed the turn to another language’s script. For an agent that turn is unusable: intent detection, entity capture and the reply all fail with it. Teams worked around it by pinning the language, which means a language menu before the conversation starts.
In 3.6 Pro the share of sessions with any wrong-script output falls from 0.55% to 0.14%, roughly one session in seven hundred. Code-switching errors fall 16% on average across the 20 public sets that mix English with German, Spanish, French, Italian or Portuguese, 19 of them improve, and none regress. You can still pin the language. Most agents can leave it on automatic.
| Error rate | Universal-3.5 Pro | Universal-3.6 Pro |
|---|---|---|
| Sessions with any wrong-script output | 0.55% | 0.14% |
Across the 17 previously supported non-English languages, error rates fall 16% on average, and every one of the 17 improves. The two public sets can disagree on a language: Common Voice error falls in 16 of the 17, by 20% on average, and FLEURS error in all of them, by 11%. Hindi is the split case, rising 19% on Common Voice while falling 39% on FLEURS.
Common Voice and FLEURS, 3.6 Pro vs 3.5 Pro
Error rate reduction by language
Higher is better
Fourteen languages are new:
- Korean
- Catalan
- Galician
- Russian
- Romanian
- Persian
- Estonian
- Afrikaans
- Marathi
- Zulu
- Urdu
- Xhosa
- Cantonese
- Norwegian Nynorsk
All 32 run through the same endpoint with automatic language detection; pass language_codes
if your app supports a fixed set.
Turns end when the caller is done, at the same latency
Every agent team owns a silence-window setting, and most have widened it once, after a caller got cut off mid-phone-number, and left it there. That is a latency cost on every turn to protect one turn in ten.
End-of-turn precision rises from 74.1% to 77.9% at held recall (92.4% to 92.8%). Median endpoint latency is unchanged at 537 ms. Both models run the identical endpointer, so the difference is the recognition model producing fewer premature end-of-turn decisions, which means a silence window you widened on 3.5 Pro can move back toward the default.
| Endpointing | Universal-3.5 Pro | Universal-3.6 Pro |
|---|---|---|
| End-of-turn precision | 74.1% | 77.9% |
| End-of-turn recall | 92.4% | 92.8% |
| Median endpoint latency | 537 ms | 537 ms |
The adaptive endpointer holds a turn open only when the transcript looks like an entity in progress: a number, a code, an email, an address. Every other turn ends the moment the caller is done, which is why the median does not move. In an ablation with the hold switched off, phone-number entity error more than triples (2.2% to 7.0%), codes and identifiers rise 29% and email addresses 17%. You are not paying latency for the hold; you are paying it on the turns where someone is still reading out a number.
Every turn also carries end_of_turn_confidence, a 0-to-1 signal. Threshold it to start your
LLM call early and discard the result if the caller keeps talking.
Speaker labels that improve during the session
Two things are new for streams with speaker labels on. The internal speaker model keeps refining as a
session runs, so live labels on new turns get better as a meeting goes on. And you can now receive
speaker revisions during the session rather than only at the end: set
speaker_labels_revision_interval_ms alongside speaker_labels=true and an
incremental revision, listing only the turns whose labels changed, arrives on that cadence. The
re-clustering runs off the transcription path, so live turn latency is unaffected. The minimum interval
is 120 seconds.
Where it stands in the field
Our own numbers are above. Two independent benchmarks run 3.6 Pro alongside everyone else’s streaming models on audio we don’t choose, which is the harder test and the one worth checking before you migrate anything.
Pipecat
Pipecat publishes an open benchmark of streaming speech-to-text services, measuring time to first speech against semantic word error rate and plotting the Pareto frontier: the set of models where nothing else is both faster and more accurate. Every model on the chart runs the same audio through the same harness, and the code and results are public.
On that run, Universal-3.6 Pro Realtime sits on the frontier: 0.96% pooled semantic word error rate at 307 ms median time to first speech, with 87.0% of transcripts exactly right. One model on the chart transcribes more accurately, Meta’s, and it sits further out on latency. That trade is the whole point of a frontier, and it is the one an agent has to make on every turn.
Coval
Coval runs an independent, continuously updated leaderboard of streaming speech models on real voice-agent audio, rerunning throughout the day rather than publishing a single dated result. Its 7-day window is the one to watch: long enough to smooth out a bad afternoon, short enough to reflect what is actually deployed right now.
Lowest word error rate
2.3%
Universal-3.6 Pro AssemblyAI
The two boards disagree on the absolute number, and they are supposed to: they run different audio through different scoring. What travels between them is the placement, which is the part worth acting on.
The numbers in this post are the headline cuts. The research post covers the training data, the evaluation sets and the methodology behind each one. The benchmarks page carries the head-to-head comparisons against other streaming models as those runs land, so it is the place to check how 3.6 Pro stacks up on the audio profile your agent actually handles.
How it works
If you are on Universal-3.5 Pro Realtime, the change is the model name. Everything else in your integration stays the same.
wss://streaming.assemblyai.com/v3/ws?model=universal-3-6-pro To receive speaker revisions during the session:
wss://streaming.assemblyai.com/v3/ws?model=universal-3-6-pro&speaker_labels=true&speaker_labels_revision_interval_ms=300000
on SpeakerRevision(msg):
for item in msg.revisions: # only the turns whose labels changed
transcript[item.turn_order].speakers = item.speakers
Test in your default settings first so you see the out-of-box behavior before tuning. The migration
guide covers language pinning, voice focus thresholds, end_of_turn_confidence and the
speaker revision parameters.
Pricing
$0.45/hr of audio, unchanged from 3.5 Pro. Volume discounts apply at any tier.
Get started
Universal-3.6 Pro Realtime is available now with your existing API key. Swap the model name, run your own calls through it, and compare the short answers, the background, and the turn boundaries against what you see on 3.5 Pro today.
Frequently asked questions
How do I switch from 3.5 Pro?
Change the model name on the connection and deploy. There is no new endpoint and nothing else in the request to migrate.
wss://streaming.assemblyai.com/v3/ws?model=universal-3-6-pro Every parameter you use today works the same way. Test in your default settings first so you see the out-of-box behavior, then revisit anything you tuned to work around 3.5 Pro, such as a widened silence window or a pinned language. Universal-3.5 Pro Realtime stays available, so you can run both against the same calls before you cut over.
Should I keep voice focus on?
It is off by default, and it should stay off unless a second person is talking near the microphone: shared offices, households, call centers with open floors. Where the challenge is noise rather than other talkers, 3.6 Pro handles it on its own.
Do I still need to pin the language?
Most agents can leave it on automatic. Wrong-script sessions fall to roughly one in seven hundred. If your app supports a fixed set of languages, passing them is still supported.
Can I move my silence window back?
If you widened it on 3.5 Pro to avoid cutting callers off, yes. End-of-turn precision is up 3.8 points at the same recall and the same median latency, and the entity hold keeps phone numbers and codes in one turn regardless of the window.
Does latency change?
Median endpoint latency is 537 ms on both models. Both were measured on matched hardware at the same operating point.
Is Universal-3.5 Pro Realtime still available?
Yes.
Where can I read how this was measured?
The research post goes through the training data, the evaluation sets and the methodology behind every number in this post. The benchmarks page carries the head-to-head comparisons against other providers and updates as new runs land.