New Universal-3.5 Pro is here. Learn more: Async Realtime
Speech-to-text evals

Is your WER benchmark lying to you?

Word error rate is easy to game and easy to misread. Better truth files, smarter normalization, and the right metric show what a model actually gets right—then see how Universal-3 holds up.

Better truth files

Sloppy reference transcripts inflate error rates. Cleaner, verified truth files reveal what a model actually gets wrong.

Smarter normalization

Casing, numbers, and punctuation shouldn't count as errors. Normalize before scoring for a fair, apples-to-apples comparison.

Measure what matters

WER misses entities, formatting, and readability. Track missed-entity rate and the metrics your users actually feel.

Metaview
Ashby
Cluely
Genio
Siro

36%

improvement in close rate

See case study
LiveKit
Earmark
Commure
Dovetail
Fireflies

“The new Universal-3.5 Pro speech model from AssemblyAI is best so far in terms of accuracy, latency, and language switching.”

Retell
CallRail
Apollo.io
ClickUp
Calabrio

80%

increase in customer satisfaction

HeyGen
Granola
Siro
JotPsych
Granola

“Assembly has saved us countless hours managing models, and provided exceptional accuracy.”

Metaview
Ashby
Cluely
Genio
Siro

36%

improvement in close rate

See case study
LiveKit
Earmark
Commure
Dovetail
Fireflies

“The new Universal-3.5 Pro speech model from AssemblyAI is best so far in terms of accuracy, latency, and language switching.”

Retell
CallRail
Apollo.io
ClickUp
Calabrio

80%

increase in customer satisfaction

HeyGen
Granola
Siro
JotPsych
Granola

“Assembly has saved us countless hours managing models, and provided exceptional accuracy.”

Quickstart

Test our models on your audio in minutes

Create a free account and run AssemblyAI's models against your own files. Compare results in the no-code playground, then wire the API into your eval pipeline.

Start building free

No credit card required

Word error rate

WER is calculated as (substitutions + insertions + deletions) / total words in the reference transcript. It's the standard metric for evaluating speech-to-text accuracy.

Average normalized WER across 26 real-world datasets.

*Lower is better*
AssemblyAI Universal-3 Pro
7.18%
ElevenLabs Scribe V2
7.65%
Azure Batch
8.56%
GPT-4o Transcribe
8.94%
Deepgram Nova 3
9.03%

Source: AssemblyAI published benchmarks, March 2026.