New Universal-3.5 Pro is here. Learn more: Async Realtime
Speech-to-text evals

Speech-to-text evals that actually understand model performance

Your WER benchmark might be lying to you. Models that look great in demos fail on real audio—noise, numbers, domain terms, overlap, and edge cases. Measure what actually breaks your pipeline.

Pick the right metric

Go beyond WER with domain-specific evals like Missed Entity Rate and Semantic WER, so the score reflects your use case.

Not all errors are equal

WER scores "gonna → going to" the same as "lisinopril → listening a pill." One keeps the meaning; one breaks your pipeline.

Tools to fix your benchmarks

Truth File Corrector, semantic word lists, and a benchmarking SDK clean up ground truth and compare models on your own audio.

Metaview
Ashby
Cluely
Genio
Siro

36%

improvement in close rate

See case study
LiveKit
Earmark
Commure
Dovetail
Fireflies

“The new Universal-3.5 Pro speech model from AssemblyAI is best so far in terms of accuracy, latency, and language switching.”

Retell
CallRail
Apollo.io
ClickUp
Calabrio

80%

increase in customer satisfaction

HeyGen
Granola
Siro
JotPsych
Granola

“Assembly has saved us countless hours managing models, and provided exceptional accuracy.”

Metaview
Ashby
Cluely
Genio
Siro

36%

improvement in close rate

See case study
LiveKit
Earmark
Commure
Dovetail
Fireflies

“The new Universal-3.5 Pro speech model from AssemblyAI is best so far in terms of accuracy, latency, and language switching.”

Retell
CallRail
Apollo.io
ClickUp
Calabrio

80%

increase in customer satisfaction

HeyGen
Granola
Siro
JotPsych
Granola

“Assembly has saved us countless hours managing models, and provided exceptional accuracy.”

Quickstart

Test with your own audio in under 10 minutes

Create a free account, drop your audio into the playground, and compare models side by side. Use the benchmarking SDK to run WER, Missed Entity Rate, and Semantic WER against your own ground truth.

Start building free

No credit card required

Word error rate

Word error rate is the standard measure of transcription accuracy, but it only tells the truth when you run it on real, varied audio—not a clean demo clip.

Pre-recorded word error rate, public benchmark datasets.

*Lower is better*
AssemblyAI Universal-3 Pro
4.50%
Mistral Voxtral Mini
5.24%
OpenAI GPT-4o Transcribe
5.34%
Deepgram Nova-3
6.66%
Azure Batch
7.02%

Source: AssemblyAI published benchmarks — assemblyai.com/benchmarks.