New Universal-3.6 Pro Realtime is now available Learn more
Insights & Use Cases

The voice agent accuracy problem nobody benchmarks

Word error rate doesn't predict voice agent quality. See how speech recognition errors compound across turns — and why entity accuracy is the metric.

Abstract green mobius illustration

Written by

Devon Malloy

Published on

30 September 2026

A patient calls to refill a prescription. She says the number clearly, at a normal pace, in a quiet room: RX-7704132.

The agent transcribes RX-7704182.

One digit. In a 19-word utterance, that's a 5% error on a single token — a rounding error on any standard benchmark, the kind of miss that would not move a word error rate figure enough to notice. And the refill fails. The lookup returns nothing, the agent apologizes, the patient repeats herself, and if she's unlucky the second attempt lands on a different digit.

This is the gap between the metric the industry publishes and the thing you actually bought.

Why standard benchmarks miss the point

Word error rate treats every word as equally important. That's what makes it easy to compute and useless for predicting whether a voice agent works.

Take the sentence above: "I need to refill my prescription, the number is RX-7704132, it's Metoprolol, 80 milligrams." A 5% WER on twenty words means roughly one word wrong. If that word is the, nothing happens. If it's Metoprolol, a pharmacist gets a call. If it's a digit in the RX number, the transaction fails outright.

Same WER. Three completely different outcomes.

The words that carry the transaction — names, account numbers, alphanumeric IDs, dates, dosages, addresses — are a small minority of tokens in any conversation and close to all of its value. They're also the hardest words to get right, because they're exactly the words a language model can't guess from context. A model can recover "I need to refuel my prescription" into "refill" from priors. It cannot recover a digit it never heard.

So a model can post an excellent WER by nailing the function words and still miss one entity in four. We've made the longer version of this argument elsewhere, but the short form is: missed entity rate is the number that predicts whether your agent completes the task. WER is context, not the frame.

Our own research says builders already know this. In AssemblyAI's Voice Agent Report, a survey of 455 voice agent builders run across Q4 2025 and Q1 2026, 76% named speech-to-text accuracy as a non-negotiable requirement — ahead of latency, ahead of cost, ahead of model sophistication. And on the other side of the call, 55% of voice agent users said the experience they hate most is being asked to repeat themselves.

Being asked to repeat yourself is what a missed entity feels like from the outside.

What the numbers actually look like

Most published speech benchmarks run on read-aloud corpora — audiobooks, scripted news, clean single-speaker recordings. Voice agents don't hear any of that. They hear interruptions, background noise, half-finished sentences, and people reciting alphanumerics under stress.

AssemblyAI's English voice-agent benchmark is built on 12,460 scripted voice-agent scenarios, which makes it a close proxy for the audio your agent actually gets, and the research write-up publishes the methodology behind the table below. You don't have to take our word for the direction, either: on Pipecat's open STT benchmark of real agent conversations, which you can rerun yourself, Universal-3.6 Pro Realtime posts a 0.96% pooled semantic word error rate, and it has the lowest word error rate on Coval's independent leaderboard (2.2%).

Model Entity error rate WER
Universal-3.6 Pro Realtime 14.4% 5.19%
Deepgram Flux EN 30.1% 13.50%
ElevenLabs Scribe v2 18.5% 7.78%
Deepgram Nova-3 26.1% 8.64%

Look at the ratio between the two columns. Every model's entity error rate is more than double its WER — and the absolute number is what should stop you: nearly a third of the entities at the bottom of the table. A model that misses nearly a third of the named entities in a conversation is still shipping a WER that looks defensible on a slide.

Broken out by type, Universal-3.6 Pro Realtime misses 10.9% of names, 10.0% of codes and IDs, and 2.4% of phone numbers. Names are a hard case, and they're hard for everyone — a surname the model has never seen has no linguistic prior to fall back on. Phone numbers and account IDs are where the transaction lives, and 2.4% is the number to hold a vendor to.

Our full methodology across the pre-recorded models — 250+ hours, 80,000+ files, 26 datasets — is published in full, alongside per-dataset results rather than a single averaged figure. The same is true of the streaming model. If a vendor gives you one number and no dataset list, you're being sold a number, not a measurement.

See Voice AI In Action

Experience natural, real-time conversations that go far beyond IVR menus. Test streaming transcription speed and accuracy on your own audio.

Try playground

The compounding mechanism

Here's the part that gets left out of vendor comparisons, and it's the reason a per-turn accuracy number understates the problem badly.

A conversation isn't one transcription. It's a chain of them. And the LLM driving your agent never hears the audio — it reads the transcript. A wrong word doesn't stay a wrong word; it enters the context window as fact, shapes the next question, and the agent proceeds confidently from a premise that was never true.

So let's do the arithmetic the industry doesn't publish.

An illustrative model, not measured data. Take a conversation with several turns that each carry at least one entity that matters — a name, a date, a number. Assume each entity is captured independently at the model's per-entity rate. Using the entity error rates above, that's a per-turn capture rate of 85.6% for Universal-3.6 Pro Realtime and 69.9% for Deepgram Flux EN. Multiply across turns and you get the probability that every entity in the conversation survived intact:

Entity-bearing turns All entities correct — 85.6%/turn All entities correct — 69.9%/turn
1 85.6% 69.9%
3 62.7% 34.2%
5 46.0% 16.7%
8 28.8% 5.7%
10 21.1% 2.8%

An appointment booking that collects a name, a date, a time, a phone number, and a reason for the visit is a five-turn conversation. At 85.6% per turn it comes through clean 46.0% of the time. At 69.9% it comes through clean 16.7% of the time.

One in six. That's not a slightly worse model. That's a different product.

Now the honest caveats, because this model is deliberately unforgiving and you should know exactly where it bends:

  • It assumes independence. Real errors cluster — a hard accent or a noisy line degrades every turn, so the good case is worse than modeled and the bad case is too.
  • It assumes no recovery. Real agents confirm. If your agent reads back every entity and the caller catches errors 70% of the time, the effective per-turn rate rises to about 95.7% for the first model and 91.0% for the second. Across five turns that's 80.2% versus 62.3% — the gap narrows, and the cost moves from failed transactions to call duration. You've bought accuracy with time, and on a $4.50/hr connected-time bill you're paying for it twice.
  • Not every turn carries an entity. A ten-turn conversation might have four that matter. Use the count of entity-bearing turns, not total turns.

None of those caveats change the shape of the curve. A per-turn advantage doesn't add across a conversation — it compounds. And the compounding is why two models that look a few points apart on a benchmark can be three to seven times apart on task completion.

Here's what that looks like in three real shapes of conversation.

Lead qualification

A prospect says "Joaquin Reyes." The transcript reads "Wakeem Race." Every subsequent turn is now anchored to a person who doesn't exist. The CRM record is wrong, the follow-up email goes to a mismatched contact, and the rep who picks it up starts the call by getting the name wrong — which is the one mistake a prospect always notices.

The agent never flagged anything. From its perspective the conversation went perfectly.

Appointment booking

A customer says "Wednesday at two." The transcript reads "Wednesday at noon." It's a common phonetic near-miss and the resulting sentence is completely plausible, so nothing downstream throws an error. The customer gets a confirmation, doesn't read it closely, and shows up two hours late.

Resolving it costs two more customer contacts and a scheduling conflict for whoever held the noon slot. The transcription error cost was one word; the operational cost was an afternoon.

Clinical intake

A patient states an allergy to Lisinopril. The transcript reads Bisoprolol. Both are cardiovascular medications, both are plausible in the context, and the substitution is phonetically reasonable — which is precisely what makes it dangerous. There's no signal anywhere in the pipeline that anything went wrong.

This is the category where the compounding math stops being an interesting spreadsheet and starts being a clinical safety question.

Build Your Voice Agent Faster

Evaluate real-time speech-to-text with low latency and strong accuracy. Launch pilots quickly with clear docs and developer-friendly APIs.

Sign up free

The specific words that break most models

Missed entities aren't randomly distributed. They fall into five categories, and knowing which one is hurting you tells you what to do about it.

Proper nouns and names. One of the hardest categories, and the numbers show it — 10.9% missed on AssemblyAI's voice-agent benchmark, against 2.4% for phone numbers. A surname outside the model's training distribution has no prior to fall back on, and the model will confidently substitute a common word that sounds similar.

Alphanumeric strings. Account IDs, prescription numbers, confirmation codes, policy numbers. These have no linguistic structure at all, so context can't rescue them — and they're usually the entity the whole call exists to capture. At 2.4% on phone numbers, roughly one call in forty needs a repeat.

Domain terminology. Drug names, part numbers, legal terms, internal product names. Adjacent to the proper-noun problem but fixable, because you know your vocabulary in advance.

Accented and multilingual speech. Both accented English and genuine code-switching mid-sentence. Most models handle a language switch at a turn boundary and fall apart when it happens inside one clause — which is how bilingual speakers actually talk.

Backchannels and disfluencies. "Mhm," "right," "uh" — these rarely change the transcript's meaning but they routinely break turn detection, causing the agent to interrupt. An interruption reads as an accuracy failure to the caller even when every word was transcribed correctly.

What "purpose-built" actually means here

That word gets used loosely. Here's what it means concretely, in parameters you can set today.

Promptability. Tell the model what domain it's operating in and it adjusts its priors for domain vocabulary before a single word is spoken. On 20,000 real voice-agent calls, detailed context cut medical-term entity errors 43%, and scenario-level context cut them 24% — the full breakdown is in the prompting and keyterms documentation.

Keyterms. Boost specific strings — medication names, account ID formats, product names — so the model recognizes them reliably. Up to 100 terms on streaming, 50 characters each, and updatable mid-stream, so you can push a caller's name into the recognizer the moment your CRM identifies them. Included at no extra cost on Universal-3.6 Pro Realtime.

agent_context. Pass your agent's own question along with the audio, and the model hears the answer through the lens of what was asked. Ask for a date of birth and it biases toward dates. Across a 20,000-file voice-agent benchmark, this alone cut WER 10.2%.

Context Carryover. Rolling conversation memory that applies earlier turns to later transcription, on by default. This is the direct counter to the compounding problem — the model gets better at a name the second time it hears it, rather than making the same mistake five turns in a row.

Turn detection you can tune. Defaults come from the mode preset you pick: min_turn_silence 128 ms and max_turn_silence 1280 ms on balanced, with a vad_threshold of 0.2. That distinguishes a mid-sentence pause from an end-of-turn signal, and it's the difference between an agent that listens and one that talks over people.

32 languages with mid-sentence code-switching, including Hinglish — intra-utterance, not per-turn. The parameter is speech_model: "universal-3-6-pro" on wss://streaming.assemblyai.com/v3/ws, with language_codes for biasing. Setup is in the streaming model selection docs.

For clinical workloads specifically, Medical Mode activates with domain: "medical-v1" and posts a 3.2% missed entity rate on medical entities — the lowest across benchmarked providers including Deepgram, Speechmatics, AWS, and Google. It's a $0.15/hr add-on, available in English, Spanish, German, and French on both pre-recorded and streaming. Details on the medical transcription page.

Teams evaluating on their own traffic land in the same place. As Foysal Osmany, Software Engineer at Fireflies, put it:

"We were searching for the best realtime ASR model for our voice agent pipeline in Fireflies. The new Universal 3.5 Pro speech model from Assembly is best so far in terms of accuracy, latency and language switching."

The right question to ask

Stop asking vendors "what's your word error rate?"

Ask: what's your missed entity rate on the words my users will actually say? Then ask which dataset, how many hours, and whether you can see the per-dataset breakdown. A vendor who leads with a single averaged number across scripted audio is answering a question you didn't ask.

Better still, don't ask at all — measure. Pull 200 of your own calls where an entity mattered, run them through both models, and count the misses by hand. It takes an afternoon and it settles the argument permanently. Our guide to evaluating speech recognition models walks through the setup, and entity accuracy in speech-to-text covers how to score entities specifically.

And here's the part the benchmark tables can't show you. The compounding math cuts both ways. If a per-turn advantage multiplies into a task-completion gap, then so does every intervention that raises the per-turn rate — a keyterm list, a scenario prompt, a confirmation step in the right place. Those aren't small optimizations that shave a percent off a metric. Applied to a five-turn conversation, they're the difference between an agent that finishes the job and one that hands the caller to a human.

Which means the most valuable thing you can know about your voice agent isn't its accuracy. It's how many entity-bearing turns your average conversation actually has. Most teams have never counted.

Unlock Voice AI ROI

Learn how automation reduces costs, shortens handle times, and scales support without adding headcount. Get guidance tailored to your industry and goals.

Talk to AI expert

Frequently asked questions

How accurate is AssemblyAI's speech-to-text?

On AssemblyAI's English voice-agent benchmark, Universal-3.6 Pro Realtime posts a 14.4% entity error rate and 5.19% WER, with 2.4% missed on phone numbers and 10.0% on codes and IDs; on Pipecat's open STT benchmark of real agent conversations it posts a 0.96% pooled semantic WER. For pre-recorded audio, English mean WER is 5.6% with a median of 4.9% across a methodology of 250+ hours, 80,000+ files, and 26 datasets. On the Hugging Face Open ASR Leaderboard we sit at 5.03 average WER, ranked #2 for private and scripted audio. Per-dataset results are on the benchmarks page.

AssemblyAI vs Speechmatics: which has better accuracy?

On the metric that predicts voice agent success — missed entity rate — AssemblyAI posts the lowest rate across benchmarked providers including Speechmatics, Deepgram, AWS, and Google, and Medical Mode reaches a 3.2% missed entity rate on medical entities specifically. Aggregate WER between leading providers is close enough that it rarely decides anything; the divergence shows up on names, alphanumeric IDs, and domain vocabulary. Run both on 200 of your own calls and count entity misses by hand rather than trusting either vendor's average.

AssemblyAI vs Google Cloud Speech-to-Text: accuracy and pricing

On AssemblyAI's English voice-agent benchmark, Universal-3.6 Pro Realtime ranks first of 21 systems on normalized WER at 5.19%, and publishes an entity error rate alongside it — 14.4% — which Google does not. On pricing, AssemblyAI streaming is $0.45/hr and pre-recorded is $0.21/hr on the flagship, billed per second with no minimums, no upfront commits, and unlimited concurrency. Current rates are on the pricing page.

AssemblyAI vs Microsoft Azure Speech-to-Text: accuracy comparison

We publish our full methodology rather than a single head-to-head number: 250+ hours across 80,000+ files and 26 datasets, with per-dataset results on the speech-to-text product page and the benchmarks hub, plus a 5.03 average WER and #2 placement for private and scripted audio on the Open ASR Leaderboard. For a voice agent decision, compare missed entity rate on your own audio rather than either vendor's averaged WER — the two providers will look far closer on WER than they do on names and account numbers.

Most accurate API for transcribing YouTube and video content

For pre-recorded video, Universal-3.5 Pro at $0.21/hr is the flagship: 18 languages with native code-switching, automatic fallback to Universal-2 for 99 languages total, our most accurate speaker diarization, and a hallucination rate roughly 30% lower than Whisper — which matters on long-form video, where hallucinated content is worse than a missed word. Benchmark results include 4.22% WER on Meanwhile and 6.77% on TedLium. Speaker labels are set up in the speaker diarization docs.

Where can I see streaming STT pricing and limits?

Streaming speech-to-text with Universal-3.6 Pro Realtime is $0.45/hr ($0.0075/min), and the fully managed Voice Agent API is a flat $4.50/hr ($0.075/min) with STT, LLM, and TTS included. Every feature is included in the $4.50/hr rate — there are no per-layer add-ons, concurrency fees, or per-agent subscriptions. Streaming is billed on session duration, the Voice Agent API on connected conversation time, both per second with unlimited concurrency and no upfront commits. The free tier covers 333 hours of streaming and 185 hours of pre-recorded audio.