Handling transcript errors: Homophones, corrections and AI quality improvement
Transcription errors can change meaning fast. Learn the most common mistakes, what causes them, and how to reduce errors in human and AI transcripts today.



Every transcript has errors in it. The useful question isn't whether yours does — it's which kind, how often, and whether the specific ones you're getting break anything downstream.
A misspelled filler word in a podcast caption costs you nothing. The same model dropping the word "not" from a clinical note, or writing a customer's account number one digit wrong, is a different category of problem entirely. Same word error rate. Wildly different consequences.
What counts as a transcription error?
A transcription error is any difference between what was said and what the transcript says. That covers three mechanical operations — a word substituted for another, a word deleted, or a word inserted that nobody said — plus a fourth category most metrics ignore: the transcript is technically correct but formatted, punctuated, or attributed in a way that changes the meaning.
Worth separating from a transposition error, which is a data-entry mistake — typing 1,342 as 1,432 when copying between systems. Transcription errors originate in the speech-to-text step; transposition errors originate downstream of it. They get conflated constantly and they have completely different fixes.
The transcription errors that actually cause problems
Six patterns account for nearly everything you'll see in a production speech-to-text pipeline.
Homophones and near-homophones
"Their," "there," and "they're" sound identical. So do "to," "too," and "two." A speech model can't hear the difference because there isn't one — it infers the right spelling from surrounding context, the way you do. It gets that wrong in a predictable place: when context is thin. Short utterances, one-word answers, and sentences that trail off give the model nothing to disambiguate against.
Omissions
Dropped words are the most dangerous error class because they're the hardest to spot. A substitution reads oddly and draws your eye. A transcript missing a single "not" reads perfectly fluently and means the opposite of what the speaker said. Omissions cluster around crosstalk, low-volume speakers, and the beginnings of turns — the moment right after someone else stops talking is where words get swallowed.
Entity errors: names, places, and numbers
Proper nouns, addresses, account numbers, and spelled-out email addresses cost real money, because they're the parts of the transcript your automation actually reads. A summary can absorb a wrong word. A workflow that routes on a customer ID can't.
They fail for a specific reason: unlike ordinary vocabulary, entities aren't predictable from context. Nothing tells the model whether the caller said "Sian" or "Shawn." That's the gap contextual prompting closes.
Language hallucination: the transcript comes back in the wrong language
This one deserves more attention than it gets, because it's common and it surprises people. Strongly accented English is sometimes transcribed as an entirely different language — the output arrives fluent, confident, and in Portuguese. Sometimes it happens even when the request carries an explicit language code.
It's a real failure mode and we're not going to pretend otherwise. Accented English remains one of the hardest cases in automatic speech recognition, ours included, and we don't claim to have solved it. What reduces it in practice:
- Set language_code explicitly rather than relying on automatic detection. Detection needs roughly 15 seconds of audio to work well, and short or noisy clips are where it drifts.
- Use contextual prompting to tell the model what domain and language it's listening to. Priming beats a bare language flag on borderline audio.
- For regional spelling, language_detection_options.localization accepts values like ["en_au"] or ["en_uk"] without switching off detection entirely.
Disfluencies: the "error" that isn't one
"Um," "uh," false starts, and repeated words are not errors. They're what people say. Whether you want them depends on the use case — a caption track shouldn't have them, a verbatim legal record must.
Formatting, punctuation, and stray tags
"Let's eat, Grandma" and "Let's eat Grandma" contain identical words. One of them is dinner. Punctuation is a meaning-bearing decision the model makes, and most accuracy metrics score it as correct either way.
A related cleanup: audio event tags are now removed from transcripts by default. Markers for laughter, music, and background sound used to appear inline and had to be stripped afterward, which routinely broke downstream text parsing.
Upload a real file — a noisy call, a multi-speaker meeting, a clip with hard names in it — and read the transcript for yourself. No integration required.
Why do transcription errors happen?
Three causes explain the overwhelming majority.
The audio doesn't carry the information. Phone codecs discard frequency range. Far-field microphones pick up the room. Two people talking at once means neither signal is clean. No model recovers what the recording never captured — that's a physics problem before it's an AI problem.
The model lacks context. Speech recognition is prediction under uncertainty. Hearing an ambiguous sound, a model picks the likeliest word given everything around it. Give it relevant context and the prediction improves; give it nothing and it falls back on general-language priors, which is precisely when your product names come out wrong.
The configuration is fighting the use case. Running language detection on 8-second clips. Requesting formatted text when you need verbatim. Leaving diarization off for a four-person meeting, then wondering why the transcript reads as one long monologue.
How to prevent transcription errors
Prime the model with context before you give it a word list
This is the biggest single improvement available and it's underused. Contextual prompting lets you hand the model a description of the domain and the situation — a meeting agenda, a prior-visit clinical note, the products and competitors likely to come up — rather than a bare list of strings.
A word list tells the model what tokens exist. Context tells it what conversation it's in, which improves words that weren't on any list. In internal healthcare testing, feeding a patient's prior-visit note cut missed medical terms by 31%.
Use keyterms prompting for the terms that must be exact
For a bounded set of strings that have to come out character-perfect — SKUs, drug names, the twelve people on the account — keyterms prompting is the right tool, and it pairs with contextual prompting rather than replacing it.
One migration note: if your code still calls word_boost, move it. That parameter deprecates on September 11, 2026. Keyterms prompting is the replacement, it's an add-on on async, and it's included at no extra cost on streaming. Lists over 100 keyterms return an error rather than a warning, and terms longer than 50 characters are ignored.
Turn on diarization for anything with more than one speaker
Attribution errors are transcription errors. A perfectly spelled transcript that assigns the customer's words to the agent is wrong in the way that matters for QA and coaching. Speaker diarization on Universal-3.5 Pro produces the transcript and the speaker changes jointly, which is what lets it hold up through short turns and rapid back-and-forth.
Fix the audio before you fix the model
Cheap wins first: get the microphone closer, record each speaker on their own channel where you can, and don't re-encode audio you've already compressed. On streaming, voice_focus isolates the primary speaker — near-field for headsets and phones, far-field for rooms and drive-thrus.
Catch what remains automatically
Every word comes back with a confidence score. Route low-confidence spans to review instead of reading whole transcripts, and pass domain-specific correction — "oz epic" to "Ozempic" — through the LLM Gateway. Sampling a fixed percentage against human-corrected references each week catches drift long before a customer reports it.
Understanding WER: what the metric does and doesn't tell you
Word error rate is the sum of substitutions, deletions, and insertions divided by the number of words in the reference. It's the industry default and it's genuinely useful for tracking one model against itself over time. It's also a badly overloaded metric for comparing vendors: it weights every word equally, so a dropped "the" scores the same as a wrong dosage; it punishes verbatim output against cleaned-up reference files; and the human ground truth behind it contains errors of its own, which quietly inflates everyone's numbers.
For reference, our published benchmarks put Universal-3.5 Pro at 4.35% average normalized WER across the evaluated datasets, and Universal-3.6 Pro Realtime posts a 5.19% WER on AssemblyAI's English voice-agent benchmark. Read those as one signal among several, not as a promise about your audio — and never as a guarantee.
Two metrics tell you more about real conversations.
cpWER — concatenated minimum-permutation word error rate — scores the transcript and the speaker attribution together. A model can't win on it by transcribing words correctly and assigning them to the wrong person, which is exactly the failure DER lets slide. Universal-3.5 Pro is optimized for cpWER rather than DER:
| Model | Average cpWER (lower is better) |
|---|---|
| Universal-3.5 Pro | 30.17 |
| ElevenLabs Scribe v2 | 35.26 |
| Gladia | 36.87 |
| Deepgram Nova-3 English | 37.92 |
Missed Entity Rate counts how often names, numbers, medications, and addresses go missing or come back wrong. It maps most directly to whether your automation works, and it's the one WER hides most effectively. More on why WER benchmarks mislead and on evaluating models on your own audio — ultimately the only benchmark that decides anything.
Published numbers are a starting point. Run Universal-3.5 Pro on the audio you actually process and compare it to what you're using now.
What errors cost when they get through
The cost of a transcription error is rarely the error. It's the human time spent finding it.
Joshua Grossberg, CTO at Kapwing, put the arithmetic plainly: "If you have an hour of content, the difference between 99% accuracy and 97% accuracy, it's a lot of time for that person to review. So you could cut down their workflow from taking half an hour, taking 20 minutes, taking 15 minutes — it's huge, right?"
Two percentage points sounds like a rounding difference until you multiply it by every hour your users process. More on that arithmetic in the true cost of inaccurate transcription, and on the gap between a clean score and a usable transcript in accuracy versus quality.
The errors you should actually chase
Here's the reframe worth taking away: stop trying to lower your overall error rate and start deciding which errors you're willing to ship.
Teams that get this right end up with an uneven transcript by design — verbatim where it gets audited, cleaned up everywhere else, contextual prompting loaded for the entities their business runs on, confidence-triggered review on nothing else. Their WER isn't dramatically better than the team next door's. Their transcripts just fail in places that don't cost anything.
That's a configuration decision, not a model decision, and you can only make it after looking at which errors your own audio produces. Model prices are on the pricing page if you're costing out the options.
Contextual prompting, keyterms, diarization, and confidence scores are all there to test. Find out which errors your recordings actually produce before you configure around a guess.
Frequently asked questions
Can I customize vocabulary and spelling in AssemblyAI transcriptions?
Yes, through three mechanisms that stack. Keyterms prompting pins specific strings like product and drug names so they come back spelled exactly as you specify. Contextual prompting primes the model with domain background, which improves terms you never listed. Custom spelling in the Speech Understanding API rewrites recognized text to your house conventions after the fact.
How do speech services handle overlapping speech in transcripts?
It depends on whether the model treats transcription and speaker attribution as one problem or two. Universal-3.5 Pro produces the transcript and the speaker changes jointly, which lets it hold short turns and crosstalk together rather than collapsing them into one speaker. Overlap is still among the hardest conditions in speech recognition — separate channels per speaker remain the most reliable fix.
How do I validate transcription accuracy automatically?
Sample a fixed percentage of production transcripts each week, correct them by hand, and score the model against those references so you're tracking drift rather than guessing. Layer confidence scores on top and route low-confidence spans to review. Track cpWER and Missed Entity Rate alongside WER, since word error rate alone won't show you an attribution failure or a missing account number.
How many different voices can a transcription API distinguish at once?
Streaming diarization with revision handles up to 10 speakers, labeling them live and then sending a single re-clustered correction within about half a second of the stream ending. For pre-recorded audio you don't have to declare a speaker count in advance, though accuracy degrades as voices get more numerous and more similar. A separate channel per speaker sidesteps the problem entirely.
How do I improve multi-speaker recognition accuracy?
Record each speaker on a separate channel if you possibly can — it turns a hard machine learning problem into a solved one. Failing that, enable diarization, use a model optimized for cpWER rather than DER, and on streaming set voice_focus to far-field for rooms and kiosks. Adding participant names through contextual prompting helps too, because the model stops guessing at who's being addressed.
What's the most accurate API for transcribing YouTube and video content?
For long-form video, use pre-recorded rather than streaming models, since async models can use the full file as context instead of committing word by word. Universal-3.5 Pro averages 4.35% normalized WER on our published benchmark datasets and handles what video throws at you — music beds, multiple speakers, and varying microphone quality inside one file. Test on your own uploads first, because published averages say very little about any specific channel's audio.
Lorem ipsum dolor sit amet, consectetur adipiscing elit, sed do eiusmod tempor incididunt ut labore et dolore magna aliqua. Ut enim ad minim veniam, quis nostrud exercitation ullamco laboris nisi ut aliquip ex ea commodo consequat. Duis aute irure dolor in reprehenderit in voluptate velit esse cillum dolore eu fugiat nulla pariatur.





