New Universal-3.6 Pro Realtime is now available Learn more
Insights & Use Cases

The production ceiling: where voice agent stacks start showing their limits

The three production ceilings voice agent builders hit after shipping, from accents to compliance to noisy environments, and how to break through each one.

Abstract green cylinder illustration

Written by

Ryan Seams

Published on

30 September 2026

Most voice agent stacks work on the demo. The demo is one speaker, in English, on a headset, in a quiet room, answering a question the agent just asked. Production is none of those things. The gap between the two is where teams lose quarters.

The short answer, up front: voice agent stacks almost always break in the same three places. Language coverage — the agent handles English and falls apart the moment a caller switches mid-sentence. Deployment constraints — the model is excellent and it cannot legally or contractually run where the audio lives. Noisy audio — the drive-thru, the contact center floor, the van, the kiosk. Each of these is a ceiling rather than a bug: you do not hit it gradually, you hit it the week you turn on a new market, a new customer, or a new deployment site, and the metrics fall off a shelf.

This post walks each ceiling, what it looks like in your logs, and the specific parameter or surface that clears it. It is deliberately not a benchmark post — the cluster hub owns the numbers in full.

First, the floor you are building on

One figure sets the context for everything below. On AssemblyAI’s English voice-agent benchmark, run against 12,460 scripted voice-agent scenarios rather than read speech, Universal-3.6 Pro Realtime posts a 5.19% word error rate. Deepgram Flux EN posts 13.50% on the same benchmark. The full table — per-category entity error rates for names, phone numbers, emails, dates, and codes — lives in the research write-up, and the measured effect of agent_context lives in the hub post; you should read them there rather than here.

The reason to state it once and move on is that raw WER is the part of this problem that is closest to solved. A 5.19% word error rate on voice-agent audio is a good floor. It is not the thing that kills your rollout. What kills your rollout is that the floor was measured under conditions your deployment does not meet — and that is what the three ceilings are about.

Ceiling 1: language coverage and mid-sentence code-switching

The first ceiling arrives the week someone opens a second market. The stack was tuned in English, the new market is bilingual, and the transcript quality does not degrade smoothly — it collapses in specific, reproducible places.

Two different problems get conflated here, and separating them is most of the work.

Problem one is breadth. How many languages does the streaming model actually support? For Universal-3.6 Pro Realtime, the answer is 32: English, Spanish, French, German, Italian, Portuguese, Arabic, Danish, Dutch, Finnish, Hebrew, Hindi, Japanese, Mandarin, Norwegian, Swedish, Turkish, and Vietnamese, plus Afrikaans, Cantonese, Catalan, Estonian, Galician, Korean, Marathi, Norwegian Nynorsk, Persian, Romanian, Russian, Urdu, Xhosa, and Zulu. Universal-3.5 Pro covers the first 18 of those on the async path.

Be suspicious of a much larger number attached to a streaming product, including ours. Our own async legacy model, universal-2, covers 99+ languages at $0.15/hr, and it exists for legacy pre-recorded integrations — not for voice agents. There is no 99-language streaming option here. If long-tail language coverage on a live socket is a hard requirement for your product, that is a real constraint to design around, and the honest move is to check every vendor’s published streaming language list rather than their overall language list.

Problem two is code-switching, and it is the one that actually breaks agents. A caller says “Sí, my confirmation number is A-as-in-alpha four seven two.” A single utterance, two languages, and an entity in the middle of the switch. A stack that detects a language per session, or even per turn, has already lost — by the time it decides, the entity is gone.

Universal-3.6 Pro Realtime handles code-switching natively inside the utterance rather than by routing between per-language models. On the async code-switching benchmark across five language pairs published in the Universal-3.5 Pro launch post, Universal-3.5 Pro posts a 7.69 average normalized WER, against 8.77 for ElevenLabs Scribe v2, 9.07 for the previous-generation Universal-3 Pro, 12.22 for Deepgram Nova-3 Multilingual, and 44.58 for OpenAI GPT-4o Transcribe. That last number is not a typo, and it is a useful reminder that a model can be excellent at monolingual transcription and still be unusable the moment two languages share a sentence.

The practical configuration is small. Pass language_codes as a list when you want to bias toward a known set — useful when you know a market is Spanish and English and nothing else. Omit it entirely when you want auto-detection and free code-switching. Teams reach for a language lock reflexively because it feels safer; on bilingual traffic it is usually the thing making the transcript worse.

What this looks like in your logs

You will see it as a WER number that is fine in aggregate and terrible on a specific slice. Segment your evaluation set by whether the utterance contains more than one language, and the two populations will not look alike. If you are only tracking pooled WER across all traffic, this ceiling is invisible until support tickets find it for you.

Run Your Own Bilingual Slice

185 hours of pre-recorded and 333 hours of streaming are included on the free tier, which is enough to segment your own code-switched traffic before you commit to anything.

Sign up free

Ceiling 2: deployment constraints

The second ceiling has nothing to do with accuracy. The model is good enough; it just cannot run where the audio is.

This shows up in three shapes, and they escalate.

Data residency. The contract says EU customer audio does not leave the EU. This is the easiest of the three to clear and the one most often discovered late, usually in a security review two weeks before launch. AssemblyAI runs EU endpoints — api.eu.assemblyai.com for async and streaming.eu.assemblyai.com for streaming — at the same price as the US endpoints, with data staying in the EU. Same models, same parameters, different hostname.

Self-hosted deployment. The next step up is a customer who will not send audio to a third-party API at all, regardless of region. Defense, some healthcare, some financial services, and a long tail of enterprise security teams land here. The deployment answer is running AssemblyAI inside the customer’s own cloud account — for example, their AWS VPC — rather than calling a hosted endpoint. That is a meaningfully different procurement conversation than “we’re SOC 2,” and it is the thing that unblocks deals the hosted-only vendors have to walk away from.

Our read of the market — and we would rather state this as our assessment than as a claim about someone else’s product — is that self-hosted is where the field thins out considerably. Streaming models are expensive to package for someone else’s infrastructure, and most vendors do not do it. Before you take our word for it, read the target vendor’s own deployment documentation and ask them directly in writing; that is the only version of this answer worth putting in a procurement doc.

Regulated data handling. If the audio contains protected health information, the question is not which cloud it runs in but whether the vendor will sign the paperwork. AssemblyAI is a business associate under HIPAA and offers a standard Business Associate Addendum (BAA), which can be reviewed and signed self-serve without a sales call. PHI redaction is available across audio and transcripts, and the platform holds SOC 2 Type 2, ISO 27001:2022, and PCI DSS v4.0. The BAA FAQ and the Business Associate Addendum are both public, which is itself the useful signal — you can read the document before you talk to anyone.

We deliberately do not publish a compliance-posture comparison table for other vendors. Asserting what another company’s legal position is, in a table cell, on our own blog, is not a claim we can stand behind, and you should discount any vendor who does it to their competitors.

The deployment question to ask on day one

Ask it before you pick a model, not after: where is the audio allowed to be, and who has to sign what? Every other decision in the stack is reversible in an afternoon. This one determines which vendors are even eligible, and it is the single most common reason a working prototype never ships. Enterprise deployment options and the security overview are the fastest way to check eligibility.

Ceiling 3: noisy audio

The third ceiling is the most physical. The agent works on a headset and degrades badly the moment the microphone is three feet away, or there is a second conversation happening four feet away, or an engine is running.

There are three separate controls here, and one wrinkle that catches almost everyone: the parameter names and the defaults are not the same on the streaming speech-to-text API and the Voice Agent API. Check which surface you are on before you copy a value across.

Control Streaming speech-to-text Voice Agent API
Noise model voice_focus — near-field or far-field session.input.voice_focus — same values, near-field by default
Suppression strength voice_focus_threshold — 0.0 to 1.0, default 0.7 session.input.voice_focus_threshold — 0.0 to 1.0, default 0.85
Latency vs accuracy mode — min_latency / balanced / max_accuracy session.input.transcription_mode — same three values

The noise model takes near-field for headsets, handsets, and phone audio where the speaker is close to the microphone, and far-field for rooms, kiosks, drive-thrus, and any setup where the microphone is across a distance and picking up reverberation and background speech. Note the hyphens; this is the actual feature, not a generic “noise suppression” toggle with intensity levels. Voice Focus is a $0.10/hr add-on on streaming and is included in the Voice Agent API’s flat rate.

The threshold is how hard that model works, on a 0.0 to 1.0 scale, and it requires the noise model to be set. Higher values suppress more aggressively. This is the dial that resolves the tradeoff every noisy site runs into: suppress too little and background speech lands in the transcript, suppress too much and you clip the quiet speaker you wanted. Start at whichever default applies to your surface — 0.7 on streaming, 0.85 on the Voice Agent API — and tune against recordings from the actual location rather than from intuition.

Both are set at connect. A mid-session change is accepted, but it only takes effect on the next speech-to-text reconnect — so in practice you choose the noise model when you open the connection and you do not get to adjust it turn by turn. For a lobby kiosk that is quiet at 7am and loud at noon, that means either accepting a setting that holds across the day or cycling the session when conditions change. For a drive-thru, where conditions vary by vehicle rather than by hour, it means picking a setting that holds for the worst realistic case rather than the average one.

The latency/accuracy control is independent of the noise settings, with the same three values on both surfaces: min_latency, balanced (the default), and max_accuracy. Unlike the noise settings it is freely mutable mid-session — so you can run a call on balanced and switch to max_accuracy for the stretch where the customer reads out a confirmation number, then switch back. Adapt accuracy on the fly; pick the noise model once, correctly.

The combination matters more than any setting alone. far-field at a raised threshold plus max_accuracy is the configuration for a noisy fixed-location deployment. near-field at the default plus balanced is the configuration for telephony, which is most contact center traffic. Setting the mode and leaving the noise model alone is the common mistake, because the mode is the parameter people find first — and the default noise model is near-field, which is exactly wrong for a drive-thru.

Noise does not degrade all output equally. Word error rate drifts up a little; entity error rate jumps, because names, street addresses, order numbers, and phone numbers have no linguistic context to fall back on when the acoustics are bad. That is why the entity columns in the voice-agent benchmark results are the ones worth reading, and why context helps most in exactly these conditions — passing the agent’s own reply into the transcription request so a short or mumbled answer resolves against what was asked.

The parameter differs by surface again here. On the streaming API it is agent_context, and that is where the published measurement comes from: across a benchmark of 10,000+ voice agent audio files, agent_context cut word error rate by 8.9%, rising to 16.4% with a context prompt on top. On the Voice Agent API the equivalent levers are session.input.transcription_prompt (up to 1750 characters) and session.input.keyterms (up to 100 terms) — agent_context is not a Voice Agent field. The rest of that breakdown is in the hub post and in the prompting and keyterms documentation.

The ceiling nobody plans for: short utterances and command-style input

There is a fourth wall that is not a ceiling in the same sense — it is a surface mismatch, and it is arguably the most common one. A large share of “voice agent” traffic is not conversation at all. It is a two-word answer. A menu selection. A push-to-talk command. A dictated field. An IVR routing decision.

Running that through a full streaming session is the wrong shape. You pay for session setup, turn detection, and end-of-turn logic to transcribe 900 milliseconds of audio where the user said “the second one.” Turn detection is genuinely good — it decides from what has been said rather than from silence alone — but none of that machinery helps when there is no turn to detect.

Two surfaces fit better.

The Sync API is a single request that returns a finished transcript, at ~134 ms p50 on a two-second clip measured from request to finished transcript, against 5–6 seconds on an async submit-and-poll flow. It takes clips from 80 ms to 2 minutes, WAV or raw PCM, across 19 languages, at $0.45/hr with keyterms prompting and conversation context included. It is built for exactly this: dictation, push-to-talk, voice commands, IVR routing, and voicemail. It does not support speaker diarization, PII redaction, or Speech Understanding, which is the correct trade for an 800-millisecond command. Product page · docs.

The Dictation API is the right call when the text goes straight into something the user is writing. It runs speech-to-text and an LLM cleanup pass in one call at $0.62/hr flat, returning finished text in about 0.36 s — resolving self-corrections to what the speaker landed on and dropping filler while keeping tone. The framing we use internally is simple: reach for Sync when you want the transcript itself, and Dictation when you want what the speaker meant to send.

If a meaningful fraction of your traffic is short utterances, splitting it off the streaming path is usually a latency win and a cost win at the same time. It is also the change least likely to be on your roadmap, because the streaming session already “works.”

Test The Split Before You Rewrite Routing

Drop a handful of your own short clips into the playground and see whether splitting command traffic off the streaming path is worth it.

Try playground

What to check before you ship

Ceiling Symptom in production The lever
Language coverage Aggregate WER fine, bilingual slice collapses 32-language streaming model with native code-switching; omit language_codes for auto-detect, or pass a list to bias
Deployment Security review blocks launch; audio cannot leave a region or an account EU endpoints, or self-hosted into the customer’s own cloud account; BAA where PHI is involved
Noisy audio Entity error rate spikes while WER drifts Noise model set to near-field or far-field, with the threshold tuned from its default (0.7 streaming, 0.85 Voice Agent); raise to max_accuracy for entity-dense stretches; pass context on every request — agent_context on streaming (−8.9% WER, −16.4% with a prompt), transcription_prompt plus keyterms on the Voice Agent API
Short utterances Session overhead dominates a one-second answer Route commands to the Sync API (~134 ms p50); route dictated text to the Dictation API (0.36 s)

One more thing worth checking that is not a ceiling at all: pricing shape. The Voice Agent API is a flat $4.50/hr ($0.075/min) with every feature included — no per-layer add-ons, no concurrency fees, no per-agent subscriptions — over a single WebSocket at wss://agents.assemblyai.com/v1/ws, with roughly one second of end-to-end latency. Stacks assembled from separate STT, LLM, and TTS vendors tend to hit a fourth ceiling that is purely commercial, and it arrives at scale rather than at launch. See pricing for the full breakdown.

Weighing A Build Against A Bundled Path

The deployment and residency questions in particular are faster to resolve in a conversation than in a docs crawl. Bring your region, your customer’s security review, and your traffic shape.

Talk to AI expert

Frequently asked questions

How many languages does AssemblyAI’s streaming model support?

Thirty-two. Universal-3.6 Pro Realtime supports English, Spanish, French, German, Italian, Portuguese, Arabic, Danish, Dutch, Finnish, Hebrew, Hindi, Japanese, Mandarin, Norwegian, Swedish, Turkish, and Vietnamese, plus Afrikaans, Cantonese, Catalan, Estonian, Galician, Korean, Marathi, Norwegian Nynorsk, Persian, Romanian, Russian, Urdu, Xhosa, and Zulu, with native code-switching inside a single utterance. There is no 99-language streaming option — the 99+ language figure belongs to universal-2, a previous-generation async model kept for legacy pre-recorded integrations.

What is the difference between voice_focus and the latency/accuracy mode?

They solve different problems and you set both. voice_focus handles the acoustics and takes near-field (the default) for headsets and phone audio or far-field for rooms, kiosks, and drive-thrus, with voice_focus_threshold controlling strength from 0.0 to 1.0 — higher is more aggressive. The mode — min_latency, balanced, or max_accuracy — trades response time against transcript quality. Watch the surface: on streaming speech-to-text the key is mode and the threshold defaults to 0.7, while on the Voice Agent API it is session.input.transcription_mode and the threshold defaults to 0.85. The noise settings are set at connect on both surfaces and only take effect on a reconnect if changed mid-session, while the transcription mode can be changed freely mid-call.

Can AssemblyAI be deployed inside our own infrastructure?

Yes. Alongside the hosted Voice AI Cloud, AssemblyAI supports self-hosted deployment into a customer’s own cloud account — for example, an AWS VPC — and EU data residency through api.eu.assemblyai.com and streaming.eu.assemblyai.com at the same price, with data staying in the EU. Which option you need is usually decided by your customer’s security review, so establish it before you pick a model.

Will AssemblyAI sign a BAA for healthcare audio?

Yes. AssemblyAI is a business associate under HIPAA and offers a standard Business Associate Addendum (BAA), which can be reviewed and signed self-serve without a sales call. AssemblyAI enables covered entities and their business associates subject to HIPAA to use the AssemblyAI services to process protected health information (PHI). PHI redaction is available across audio and transcripts, and the platform holds SOC 2 Type 2, ISO 27001:2022, and PCI DSS v4.0.

When should I use the Sync API instead of a streaming session?

Use the Sync API when the audio is a short, bounded clip rather than a conversation — voice commands, push-to-talk, IVR routing, voicemail, dictated fields. It returns a finished transcript in a single request at roughly 134 ms p50 for a two-second clip, handles audio from 80 ms to 2 minutes across 19 languages, and costs $0.45/hr. Streaming is the right surface when there are turns to detect; Sync is the right surface when there are not.

How does Universal-3.6 Pro Realtime compare to Deepgram Flux on real agent audio?

On AssemblyAI’s English voice-agent benchmark, which uses 12,460 scripted voice-agent scenarios rather than read speech, Universal-3.6 Pro Realtime posts a 5.19% word error rate and a 14.4% entity error rate, against 13.50% and 30.1% for Deepgram Flux EN. The full per-category breakdown — names, phone numbers, emails, dates, codes — is published in our research write-up and on the benchmarks page.