
Speech-to-text accuracy is usually reported as word error rate—the share of words a model substitutes, deletes, or inserts against a human reference. The number vendors publish comes from clean read-speech benchmarks, and it does not predict what you'll see in production. On LibriSpeech Clean, Universal-3 Pro records 1.52% WER. On AssemblyAI's English voice-agent benchmark of 12,460 scripted voice-agent scenarios, Universal-3.6 Pro Realtime records 5.19%. Different models, different modes, so don't read those as a before-and-after—but the second number is the one your agent actually has to survive.
That gap is the whole subject of this post.
Here's why it matters more for voice agents than for anything else built on speech. A transcription tool with a bad transcript produces a document someone can fix. An agent with a bad transcript takes an action. It books the wrong date, retrieves the wrong account, confirms a phone number that doesn't exist, and does all of it confidently, because nothing downstream of the transcript has any way to know the transcript was wrong.
Builders agree on this, at least in the abstract. In AssemblyAI's 2026 Voice Agent Report—a survey of 455 practitioners fielded across Q4 2025 and Q1 2026—76% chose speech-to-text accuracy as a non-negotiable capability, and 52.5% named accuracy and misunderstandings as their single biggest challenge in production. Another 45% report frequently misheard words, and 47.5% are targeting better than 95% accuracy.
Which sounds like a settled question until you ask any of them which number they're optimizing—and discover most are watching a metric that can't see the errors breaking their product.
What speech recognition accuracy actually measures
Accuracy, as reported, counts correct words against total words spoken. Every word carries equal weight.
That equal weighting is the entire problem, and it's easiest to see with an example. "Patient has no allergies" becoming "patient has known allergies" is a one-word substitution. In word error rate terms it's a rounding error on a long transcript. In clinical terms it inverts the meaning of the sentence and could kill someone.
Now scale that down to something more mundane. Your agent asks for a callback number. The caller says "five-five-five, oh-one-four-two." The model returns "five five five oh one forty two." Every word is arguably correct. WER is zero. The number is wrong, the callback fails, and the metric you're watching reported a perfect score.
So there are really two questions hiding inside "how accurate is this model," and they have different answers:
- How many words did it get wrong? This is word error rate. It's cheap to compute, universally reported, and comparable across vendors.
- Did it get the words that mattered wrong? This is entity accuracy, and almost nobody publishes it.
Everything downstream of a transcript inherits its errors, and inherits them silently. That's what makes accuracy a bottleneck rather than a feature—a slow pipeline announces itself, a wrong pipeline doesn't.
Dr. Shane Lynn, CEO of EdgeTier, which builds conversational intelligence and agent evaluation on top of transcripts, puts the compounding problem better than we can:
"The transcript quality is critical, both for user perception and our AI models. Once you lose trust in transcript accuracy, you erode trust in the product. For text classification, phrase detection, and agent evaluation, the language has to be correct — otherwise, the whole system falls apart."
Dr. Shane Lynn, CEO, EdgeTier
Word error rate: what it's good for and what it hides
WER is the sum of substitutions, deletions, and insertions divided by the total words in the reference transcript, expressed as a percentage. Zero is perfect. The formula is simple, which is most of why it has survived as the industry default despite treating "um" and "amoxicillin" as equally important.
We've written a full argument for why the metric is structurally broken in word error rate is broken. Rather than repeat those here, this post takes the narrower question: what does WER hide specifically in an agent pipeline?
Three things.
It hides where the errors are. Two models with identical 8% WER can have wildly different entity error rates. One scattered its errors across filler words; the other concentrated them on proper nouns. Your agent cares enormously which one it got.
It hides the reference problem. WER is measured against a human transcript, and human transcripts contain errors, inconsistent filler-word conventions, and arbitrary number formatting. A significant chunk of any published WER figure is disagreement about the ground truth rather than model error.
It hides the audio. LibriSpeech is audiobook narration—clean, articulate, consistently paced, recorded by people who are professionally good at reading aloud. There is no version of your production traffic that resembles it.
Why is my production accuracy worse than the vendor's benchmark?
Because the benchmark measured a different problem.
Read speech in a quiet room is close to a solved task. Conversational speech over a compressed phone line, with two people occasionally talking over each other and a dog in the background, is not. When you move a model from the first condition to the second, the error rate goes up—and it goes up unevenly, concentrating on exactly the content your application depends on.
Four factors do most of the damage.
Audio quality and background noise
Signal-to-noise ratio has an outsized, non-linear effect. Small degradations in SNR produce disproportionate increases in error rate, and past a certain point recognition degrades sharply rather than gracefully. Far-field audio—a kiosk, a drive-thru, a speakerphone across a conference table—is the common production case that benchmarks essentially never cover.
Two things help on the streaming speech-to-text path. Universal-3.6 Pro Realtime's voice_focus isolates the primary speaker and suppresses background speech, with a "near-field" setting for headsets and phones and a "far-field" setting for rooms and kiosks. And its max_accuracy mode trades latency for robustness on exactly this audio.
Domain vocabulary and out-of-vocabulary words
Specialized fields carry vocabulary a general model has rarely or never seen: drug names, part numbers, internal product codenames, regional place names. An out-of-vocabulary error doesn't stay contained either—the model tries to make the surrounding words consistent with its wrong guess, so one unknown term corrupts a clause.
This is the failure mode contextual prompting exists to fix, covered in the optimization section below.
Speaker diarization accuracy
In multi-speaker audio, you have two independent ways to be wrong: the words and who said them. A transcript with perfect words and swapped speaker labels is often more damaging than one with a few wrong words, because it attributes a commitment to the wrong party.
Most diarization is benchmarked on DER—diarization error rate—which scores speaker segmentation in isolation from the transcript. Universal-3.5 Pro optimizes for cpWER instead, concatenated minimum-permutation word error rate, which scores the transcript and the attribution together. That's a meaningful choice: it means the model is being tuned on the thing users experience rather than the thing that's easy to measure.
| Model | Average cpWER (lower is better) |
|---|---|
| Universal-3.5 Pro | 30.17 |
| ElevenLabs Scribe v2 | 35.26 |
| Gladia | 36.87 |
| Deepgram Nova-3 English | 37.92 |
Measured across meetings, telephony, far-field, and conversational audio. If you want the mechanics, how speaker diarization works covers them, and the speaker labeling docs cover the parameters.
Accented, multilingual, and code-switched speech
Code-switching—a speaker moving between languages mid-sentence—is common in real conversation and almost absent from benchmarks. Handled badly, the model either transcribes the second language phonetically in the first, or drops it.
| Model | Average normalized WER, 5 language pairs (lower is better) |
|---|---|
| Universal-3.5 Pro | 7.69 |
| ElevenLabs Scribe v2 | 8.77 |
| Deepgram Nova-3 Multilingual | 12.22 |
| OpenAI GPT-4o Transcribe | 44.58 |
Coverage, since it's the immediate follow-up question: Universal-3.5 Pro covers 18 languages with native code-switching built into the model rather than stitched on as a second pass. Universal-2 remains the path for the long tail of 99+ languages.
Published benchmarks measure read speech in quiet rooms. Get a free API key and run your own noisy, accented, real-world recordings through Universal-3.5 Pro instead.
How the major speech-to-text providers compare on accuracy
The benchmark worth comparing on is the one measured on the audio you actually have. For agent pipelines that's AssemblyAI's English voice-agent benchmark, which is built from 12,460 scripted voice-agent scenarios rather than read speech. Lower is better throughout.
| Metric | Universal-3.6 Pro Realtime | Deepgram Flux EN | ElevenLabs Scribe v2 | Deepgram Nova-3 |
|---|---|---|---|---|
| Word error rate | 5.19% | 13.50% | 7.78% | 8.64% |
| Entity error rate | 14.4% | 30.1% | 18.5% | 26.1% |
| Names | 10.9% | 29.0% | 14.8% | 24.3% |
| Codes / IDs | 10.0% | 46.1% | 12.0% | 27.6% |
| Phone numbers | 2.4% | 11.5% | 3.4% | 4.5% |
Independent results point the same way: on Pipecat's open STT benchmark of real agent conversations, Universal-3.6 Pro Realtime posts a 0.96% pooled semantic WER.
Compare the first row with the second and you get the argument of this entire post in two lines of a table.
On word error rate the spread is 5.19% to 13.50%—meaningful, but the kind of gap a procurement team might reasonably decide isn't worth switching for. On entity error rate the spread is 14.4% to 30.1%. One of those numbers means nearly a third of the entities in your conversation come back wrong.
And entities are the only part your agent uses. Nobody's booking system needs the filler words.
Note the mode difference before anyone builds a slide out of this: the voice-agent benchmark figures are streaming measurements on Universal-3.6 Pro Realtime. The 1.52% LibriSpeech Clean figure quoted in the intro is Universal-3 Pro on async, on read speech. They are two separate facts about two different models in two different modes, and dividing one by the other produces a number that means nothing. Both are published with full methodology.
Some accuracy failures barely move WER at all
The clearest evidence that WER is the wrong lens is the category of failure that makes a transcript obviously worse while barely touching the score. Four structural examples, none of which a benchmark table would surface:
- Hallucinated insertions. A model that invents a short agreement token—a "yes" nobody said—adds one word against thousands. WER moves by a fraction of a percent. An agent reading that transcript sees a customer consenting to something.
- Confidence scores you can't trust. If your pipeline routes low-confidence turns to a human, the confidence score is load-bearing. WER says nothing whatsoever about whether that score is well calibrated, so a model can post an excellent error rate while making your review queue useless.
- Punctuation and casing. Both are normally stripped before WER is computed, so improvements and regressions in either are literally invisible to the metric—and immediately visible to anyone reading the output.
- Speaker attribution. WER scores the words without asking who said them. A transcript with perfect words and swapped speakers scores perfectly and is worse than useless in a compliance review.
Every one of those is a real accuracy problem, and none of them shows up in a headline number. That's the case for testing on your own audio rather than reading anyone's leaderboard, including ours—and it applies just as much to async speech-to-text workloads as to live agent traffic.
How do you test speech-to-text accuracy on your own audio?
Build a test set from 20–50 real recordings. Not clean samples, not the demo file—the accents, noise floor, vocabulary, and call length you actually get. If 30% of your traffic is far-field, 30% of your test set should be.
Then normalize before you score, or your results are garbage. Both the reference and the hypothesis need the same treatment:
Python:
import re
def normalize(text: str) -> str:
text = text.lower()
text = re.sub(r"[^\w\s']", " ", text) # drop punctuation, keep apostrophes
text = re.sub(r"\s+", " ", text)
return text.strip()Expand contractions consistently, standardize number formats (decide once whether "twenty five" or "25" is canonical, and apply it to both sides), and decide whether filler words count. Then apply the same rules to every provider you test, or you're measuring your normalizer instead of the models.
Score three things, not one:
- Word error rate, for comparability with published figures.
- Entity error rate on the entity classes your application depends on. Tag them in the reference and score them separately. This is the number that predicts whether your agent works.
- Task completion. Did the booking succeed? Did the lookup return the right record? This is the only metric your users experience, and it's the one that should break a tie.
How to evaluate speech recognition models covers the full methodology, including how to build references without paying for professional transcription.
One warning about metrics that look relevant and aren't: don't evaluate an agent pipeline on time-to-first-byte or real-time factor. TTFB measures your network, and RTF measures batch throughput. For an agent, the numbers that matter are emission latency—how long from a word being spoken to it arriving—and time-to-complete-transcript once the turn ends.
You've just been told to test on your own audio. Here's the fastest way to do it—drop a real recording into the playground and compare streaming accuracy and latency in the browser.
How to improve speech-to-text accuracy
Three levers, in descending order of how much they move the number for an agent pipeline. The first one is the one almost nobody uses.
1. Pass your agent's own question to the model
Your agent knows what it just asked. The speech model doesn't—unless you tell it.
That asymmetry is the largest recoverable accuracy loss in a typical voice agent, and it's a one-parameter fix. Pass the agent's last utterance into agent_context on the streaming session, and the model uses it to disambiguate the reply:
Python:
from assemblyai.streaming.v3 import RealTimeParameters
params = RealTimeParameters(
sample_rate=16000,
speech_model="universal-3-6-pro",
agent_context=(
"You are a pharmacy scheduling assistant. "
"You just asked: 'Which medication is this refill for?'"
),
)Note the singular speech_model—streaming uses the singular parameter, async uses the plural speech_models array. The sample targets Python SDK 1.0.0, released August 14, 2026.
Across a benchmark of 20,000 voice agent audio files, passing agent context cut word error rate by 10.2%. The breakdown is more interesting than the headline: fabrications down 18.3%, hallucinations down 17.2%, place-name entities down 15.5%, short-utterance errors down 13.7%, name entities down 9.4%, medical entities down 9.4%.
Short-utterance errors deserve their own sentence. "Yep," "nope," "that's the one," and "no, the other one" are most of what a caller says to an agent, they carry disproportionate decision weight, and they're the hardest thing in speech recognition to get right without knowing what question preceded them.
It's included in the $0.45/hr base price on Universal-3.6 Pro Realtime, along with the rolling Context Carryover memory that keeps prior turns in scope. You are not paying extra for this.
2. Prime the model with your domain vocabulary
The mechanism here has changed generation over generation, and the change is worth understanding rather than just copying the new parameter name.
The old approach was a term list: hand the model 200 or 1,000 strings and hope the right ones match. Universal-3.5 Pro supports contextual prompting instead—prime it with actual context. A meeting agenda. A prior-visit clinical note. Your product and competitor names in a sentence rather than a list.
The difference isn't cosmetic. In an internal healthcare test, feeding a patient's prior-visit note cut missed medical terms by 31%, even when the note came from an earlier visit. A term list can't do that, because the useful signal was the relationship between the terms, not the terms themselves.
Keyterm prompting still exists where you want it: Universal-3.5 Pro supports up to 1,000 terms against Universal-2's 200, and on Universal-3.6 Pro Realtime keyterm prompting is included at no extra cost. You can combine a prompt and a keyterm list in the same request—they are not mutually exclusive, and the prompting and keyterms docs show both set together on a streaming session.
3. Fix the audio, where you control it
Better input helps. It always helps. But it's third on this list for a reason: in most production deployments you don't control the microphone. The caller is on whatever phone they own, in whatever room they're standing in, and no amount of hardware guidance in your docs changes that.
Where you do control it—kiosks, drive-thrus, in-house contact center headsets—spend the money. Where you don't, voice_focus and max_accuracy mode are doing the equivalent job in software. Voice Focus is a +$0.10/hr add-on on Universal-3.6 Pro Realtime; max_accuracy costs nothing but latency.
One lever that used to appear in this section and shouldn't: custom model fine-tuning. AssemblyAI doesn't sell it, and recommending it here was recommending a product that doesn't exist. Contextual prompting is the supported path to domain adaptation, it takes minutes instead of months, and it doesn't create a model you have to retrain every time your vocabulary changes.
4. Measure the outcome, not the metric
The last lever isn't a setting. It's what you watch after shipping.
Track booking rates, resolution rates, and task completion—then A/B test model configurations against those, not against an offline WER score. A configuration that improves WER by half a point and task completion by nothing is not an improvement, it's a rounding error you spent a week on. A configuration that leaves WER flat and lifts booking completion by 4% is the one you ship, and an offline benchmark would have told you to skip it.
Final words
Here's the thing nobody says out loud about speech-to-text procurement: most teams pick a model on a number that was measured on audio they'll never encounter, then spend the next six months debugging the consequences without ever connecting the two.
The fix isn't a better benchmark. It's a different question. Stop asking "which model has the lowest WER" and start asking "which model gets the twelve entity types my product actually acts on." Those have different answers, and only one of them is published.
What's genuinely changed in the last year isn't that models got more accurate—though they did. It's that the biggest remaining accuracy gains stopped coming from the model and started coming from what you tell it. A 10.2% WER reduction from passing your agent's own question into the session is not a model improvement. It's information you already had, that you weren't sending.
Go check whether you're sending it.
For the wider argument about what accuracy means once a human is reading the transcript rather than an agent acting on it, transcription accuracy vs transcription quality takes the perception side of this.
Evaluate real-time speech-to-text with low latency and strong accuracy. Launch pilots quickly with clear docs and developer-friendly APIs.
Frequently asked questions
How to improve speech-to-text accuracy?
Three levers move the number most, in this order: pass conversational context to the model, prime it with your domain vocabulary, and fix the audio at the source. On AssemblyAI, passing the agent's own question via agent_context cut word error rate by 10.2% across 20,000 voice agent files, with fabrications down 18.3%. Contextual prompting and audio improvements compound on top of that.
Why is my talk to text so inaccurate?
Almost always background noise, domain vocabulary, or both. Signal-to-noise ratio has a non-linear effect on error rate, so a room that sounds only slightly worse to you can be substantially worse to a model, and specialized fields carry vocabulary a general model has rarely seen. Clean benchmark scores are measured on read speech in quiet rooms, which is why they do not predict what you see on a real call.
What is the most accurate speech-to-text model for voice agents?
On AssemblyAI's English voice-agent benchmark of 12,460 scripted voice-agent scenarios, AssemblyAI's Universal-3.6 Pro Realtime records a 5.19% word error rate against Deepgram Flux EN's 13.50%, ElevenLabs Scribe v2's 7.78%, and Deepgram Nova-3's 8.64%. The gap is widest on entities: 14.4% entity error rate versus Deepgram Flux EN's 30.1%, which matters because entities are what voice agents act on.
What is a good word error rate for speech-to-text?
It depends entirely on what the transcript feeds. For summaries and meeting notes, 10–15% WER is usually survivable because meaning is preserved. For a voice agent capturing account numbers or medication names, headline WER is the wrong metric—track entity error rate instead, since one wrong digit fails the task regardless of how well the surrounding words scored.
How do you test speech-to-text accuracy on your own audio?
Build a test set from 20–50 real recordings that reflect your actual conditions—the accents, noise floor, and vocabulary you get in production, not clean samples. Transcribe them with each candidate provider, normalize both sides identically before scoring, and report entity accuracy alongside WER. Testing on public benchmarks tells you how a model handles read speech, not your calls.
Can AI speech recognition match human transcription accuracy?
On clean audio it already exceeds most human transcriptionists—Universal-3 Pro records 1.52% WER on LibriSpeech Clean. On overlapping speech, heavy accents, and noisy far-field audio, trained humans still lead. The practical answer is that the gap has narrowed to specific hard conditions rather than existing across the board.
Lorem ipsum dolor sit amet, consectetur adipiscing elit, sed do eiusmod tempor incididunt ut labore et dolore magna aliqua. Ut enim ad minim veniam, quis nostrud exercitation ullamco laboris nisi ut aliquip ex ea commodo consequat. Duis aute irure dolor in reprehenderit in voluptate velit esse cillum dolore eu fugiat nulla pariatur.





