The first API built for dictation. Your users speak, and it returns text that is ready to send: filler gone, self-corrections resolved, names spelled right.
The first API built for dictation. Your users speak, and it returns text that is ready to send: filler gone, self-corrections resolved, names spelled right.
Universal-3.6 Pro Realtime measured against every major streaming vendor on the same audio.
Word error rate
WER is calculated as (substitutions + insertions + deletions) / total words in the reference transcript. It is the standard metric for evaluating speech-to-text accuracy.
Average normalized WER across selected datasets
*lower is better*
4.35%
5.24%
5.34%
5.50%
5.87%
5.91%
6.47%
6.66%
7.02%
7.16%
17.39%
AssemblyAI Universal-3.5 Pro
Mistral Voxtral Mini
OpenAI GPT-4o Transcribe
Cohere Transcribe
ElevenLabs Scribe V2
Qwen3 ASR
Gladia
Deepgram Nova-3
Azure Batch
Grok
Soniox
WER by dataset expand_more
Dataset
AssemblyAI Universal-3.5 Pro
ElevenLabs Scribe V2
Mistral Voxtral Mini
OpenAI GPT-4o Transcribe
Cohere Transcribe
Qwen3 ASR
Gladia
Deepgram Nova-3
Azure Batch
Grok
Soniox
Synthetic medical
0.33%
0.41%
1.25%
0.55%
1.33%
1.15%
1.23%
0.51%
1.55%
1.38%
0.72%
Accented English (India)
5.19%
5.92%
6.44%
6.49%
6.39%
6.61%
6.78%
7.77%
8.04%
6.60%
53.48%
General speech
6.24%
7.19%
6.64%
7.45%
8.35%
7.01%
11.08%
8.81%
8.42%
9.76%
7.55%
Webinar speech
5.63%
9.96%
6.65%
6.87%
5.91%
8.85%
6.79%
9.55%
10.07%
10.90%
7.79%
Average
4.35%
5.87%
5.24%
5.34%
5.50%
5.91%
6.47%
6.66%
7.02%
7.16%
17.39%
Missed entity rate
Missed entity rate measures errors on named entity categories including names, organizations, locations, medical terms, money, occupations, temporal expressions, and URLs. Unlike aggregate WER, it isolates the tokens that carry the most semantic weight in downstream applications.
Missed entity rate by provider
*lower is better*
Name
Organization
Location
Medical
Money
Occupation
Temporal
Url
24.1
19.0
8.7
13.9
43.7
9.0
16.3
54.4
25.9
20.8
10.0
19.8
39.2
9.4
6.2
65.0
29.2
23.4
12.2
16.1
37.3
13.0
13.8
65.5
20.9
16.6
11.5
10.8
76.9
9.1
20.4
47.1
25.6
19.3
13.6
15.7
76.8
9.7
19.4
54.8
36.1
31.6
24.3
31.1
80.4
10.6
29.7
98.1↑
AssemblyAI Universal-3.5 Pro
Mistral Voxtral Mini
OpenAI GPT-4o Transcribe
ElevenLabs Scribe V2
Deepgram Nova-3
NVIDIA Canary 1B
MER by entity type expand_more
Entity type
AssemblyAI Universal-3.5 Pro
Mistral Voxtral Mini
OpenAI GPT-4o Transcribe
ElevenLabs Scribe V2
Deepgram Nova-3
NVIDIA Canary 1B
Name
24.08%
25.89%
29.23%
20.87%
25.58%
36.08%
Organization
18.99%
20.80%
23.40%
16.57%
19.25%
31.60%
Location
8.72%
9.98%
12.21%
11.54%
13.57%
24.27%
Medical
13.87%
19.78%
16.07%
10.80%
15.69%
31.07%
Money
43.70%
39.22%
37.30%
76.88%
76.82%
80.40%
Occupation
9.01%
9.43%
13.01%
9.12%
9.67%
10.59%
Temporal
16.33%
6.21%
13.78%
20.42%
19.42%
29.70%
Url
54.37%
65.05%
65.53%
47.09%
54.79%
98.06%
Diarization
Diarization segments audio by speaker. We report cpWER (concatenated minimum-permutation word error rate), which jointly evaluates transcription and speaker assignment.
Average cpWER across DiPCo, CALLHOME, NOTSOFAR, and AMI
*lower is better*
30.17%
30.35%
35.26%
36.60%
36.88%
37.52%
37.93%
42.58%
50.64%
114.81%
AssemblyAI Universal-3.5 Pro
Azure
ElevenLabs Scribe V2
Speechmatics
Gladia
Mistral Voxtral
Deepgram
Grok
Google
Soniox
cpWER by dataset expand_more
Dataset
AssemblyAI Universal-3.5 Pro
Azure
ElevenLabs Scribe V2
Speechmatics
Gladia
Mistral Voxtral
Deepgram
Grok
Google
Soniox
DiPCo
33.48%
33.23%
48.78%
36.88%
41.90%
38.04%
38.31%
41.60%
56.40%
126.82%
CALLHOME dev/eval
17.18%
20.28%
14.82%
21.41%
20.06%
19.50%
23.70%
36.85%
29.89%
103.85%
CALLHOME train
17.78%
20.14%
16.18%
20.61%
21.02%
19.36%
25.47%
33.68%
31.12%
102.55%
NOTSOFAR test
37.02%
35.68%
44.37%
56.53%
47.81%
51.91%
48.24%
52.35%
63.83%
111.86%
NOTSOFAR dev
48.22%
45.38%
51.29%
63.38%
56.12%
61.42%
62.55%
58.92%
75.63%
116.41%
AMI
27.36%
27.39%
36.14%
20.82%
34.34%
34.86%
29.28%
32.10%
46.96%
127.36%
Average
30.17%
30.35%
35.26%
36.60%
36.88%
37.52%
37.93%
42.58%
50.64%
114.81%
Multilingual
Global WER across the evaluated multilingual benchmark suite, with language-level breakdowns where available.
Global multilingual WER
*lower is better*
7.84%
8.22%
9.52%
10.50%
11.13%
13.36%
14.39%
15.71%
AssemblyAI Universal-3.5 Pro
Speechmatics Enhanced
OpenAI GPT-4o Transcribe
Mistral Voxtral Mini
ElevenLabs Scribe V2
Cohere Transcribe
OpenAI Whisper-1
Deepgram Nova-3
Global WER and language breakdown expand_more
Model
Global WER
Evaluated languages
German
Spanish
French
Italian
Portuguese
AssemblyAI Universal-3.5 Pro
7.84%
5/20
9.07%
6.57%
10.17%
6.74%
6.65%
Speechmatics Enhanced
8.22%
5/20
10.30%
5.84%
10.09%
7.88%
7.01%
OpenAI GPT-4o Transcribe
9.52%
15/20
—
—
—
—
—
Mistral Voxtral Mini
10.50%
10/20
—
—
—
—
—
ElevenLabs Scribe V2
11.13%
19/20
—
—
—
—
—
Cohere Transcribe
13.36%
—
—
—
—
—
—
OpenAI Whisper-1
14.39%
—
—
—
—
—
—
Deepgram Nova-3
15.71%
—
—
—
—
—
—
Code switching
Code-switching benchmarks test Common Voice-derived English paired with German, Spanish, French, Italian, and Portuguese.
Average code-switching WER
*lower is better*
7.60%
10.07%
12.42%
12.83%
21.88%
22.72%
22.72%
28.13%
36.05%
39.38%
42.71%
46.19%
46.39%
51.24%
52.40%
AssemblyAI Universal-3.5 Pro
ElevenLabs Scribe V2
Soniox
Deepgram Nova-3
Qwen3 ASR
Mistral Voxtral Mini
VibeVoice
Cohere Transcribe
Speechmatics
Grok
OpenAI GPT-4o Transcribe
Gladia
AWS Transcribe
Azure Batch
Gemini
WER by language pair and condition
*lower is better*
en-es long
en-de short
en-fr short
en-it short
en-pt short
5.9
5.9
7.7
6.9
11.7
5.5
10.6
11.4
11.6
11.3
8.5
11.8
14.8
11.6
15.5
10.1
12.7
14.2
12.5
14.6
9.1
24.9
27.0
22.0
26.4
7.8
27.1
29.6
27.1
21.9
11.4
24.4
27.4
22.8
27.6
28.8
27.4
32.5
27.3
24.6
27.3
32.8
35.8
43.0
41.5
12.5
46.7
47.3
42.5
47.9
34.9
42.9
45.5
46.6
43.7
49.7
47.1
42.4
49.4
42.3
43.0
46.0
51.2
47.7
44.1
48.5
55.6↑
54.7
50.8
46.5
47.5
53.2
56.9↑
55.4↑
49.0
AssemblyAI Universal-3.5 Pro
ElevenLabs Scribe V2
Soniox
Deepgram Nova-3
Qwen3 ASR
Mistral Voxtral Mini
VibeVoice
Cohere Transcribe
Speechmatics
Grok
OpenAI
Gladia
AWS Transcribe
Azure Batch
Gemini
Price-performance
Choose a metric to see how each provider's accuracy compares at their price point.
Price per hour against the selected metric
*bottom-left is best*
Pricing from Artificial Analysis.
Demo
Listen to the difference
Play the preloaded sample and compare the transcripts side by side.
Medication name changes
Ramipril becomes a different medication, which can change the patient record.
Just use my personal email. It's jonjay@freemail.com. That's J-O-N-J-A-Y@freemail.com. So jonjay@freemail.com.
AssemblyAI
Just use my personal email. It's jonjay@freemail.com. That's J-O-N-J-A-Y@freemail.com. So jonjay@freemail.com.
Competitor
Just use my personal email. It's johnjay@female.com. That's jonjay@freemail.com. So johnj@freemail.com.
Multiple speakers collapse into one label
Distinct turns from Chris, Sally, Donald, and Jenny are merged under speaker 0.
0:00 / 0:00
import assemblyai as aai
aai.settings.api_key = "YOUR_API_KEY"
config = aai.TranscriptionConfig(
speech_model="universal-3.5-pro",
speaker_labels=True,
)
transcript = aai.Transcriber().transcribe(
"meeting.mp3", config=config
)
for u in transcript.utterances:
print(f"Speaker {u.speaker}: {u.text}")
Truth
Chris:yeah
Sally:mmm
Donald:i was thinking in about uh four months they're having the county fair perhaps you can set a booth up at the county fair
Chris:that's also a great idea um we would have to coordinate that with the city would anyone um be willing to take care of that
Jenny:yes i can take care of that and talk to the city coordinate it with them i think it's a great idea that we have it at the fair like uh donald mentioned
Chris:um ok so uh what else uh does anybody have any ideas of what
AssemblyAI
A:Yeah, I was thinking about, uh, 4 months they're having the county fair. Perhaps we could set a booth up at the county fair.
B:That's also a great idea. Um, we would have to coordinate that with the city. Would anyone be willing to take care of that?
C:Yes, I can take care of that and talk to the city, coordinate it with them. I think it's a great idea that we have it at the fair like Donald mentioned.
B:Um, okay, so, uh, what else? Uh, does anybody have any ideas of what
Competitor
Deepgram diarization
0:Yeah. Mhmm. I was thinking that in about four months, they're having the county fair. Perhaps you could set a booth up at the county fair. That's also a great idea. We would have to coordinate that with the city. Would anyone be willing to take care of that? Yes. I can take care of that and talk to the city, coordinate it with them. I think it's a great idea that we have it at the fair, like Donald mentioned.
0:Okay. So what else? Does anybody have any ideas of what
A band name changes meaning
NOFX becomes No Effect, turning a name into a phrase.
Technical reportRealtime speech-to-textdata as of 2026-09-16v2026.09
Realtime speech recognition, measured against the field
Universal-3.6 Pro Realtime and 26 competitor models on 5 evaluation suites: same audio, same scoring pipeline.
Model under test Universal-3.6 Pro Realtime (production endpoint)Evaluation AssemblyAI ResearchRivals run via each vendor’s public streaming API, default settingsSnapshot 2026-09-16
Universal-3.6 Pro Realtime has the lowest entity error rate on the voice-agent suite: 14.41% against 18.19% for the nearest competing system, a 21% lead. It also leads on code-switched speech.
Every score is in the tables. §9 reproduces independent evaluations; §10 describes the method.
At a glance
Voice-agent entity error
14.41%
1st of 22
next best Soniox Realtime 18.19%
Code-switching (5 pairs)
7.20%
1st of 23
next best AWS Transcribe 9.13%
Voice-agent WER, heavy noise
6.13%
1st of 22
next best Soniox Realtime 7.68%
English short-form WER
3.44%
1st of 22
next best Reson8 3.70%
Voice-agent WER, Indian-accented
6.64%
1st of 22
next best Soniox Realtime 7.78%
01Overview
Each cell is the vendor’s best model on that metric, coloured by rank within the column (solid blue is first). Values are in the tables of §2–§8.
Rank across 9 headline metrics, vendors ordered by mean rank.
Vendor (best model per cell)
VA entity err. NEER %
Entity (judge) %
Medical MER %
Diarization DER %
Non-speech resp. %
Code-switching WER %
Multi-lingual FLEURS %
EN short-form WER %
EN long-form WER %
meanrank
AssemblyAI
14.4
18.1
4.2
16.0
0.3
7.2
8.4
3.4
8.8
3.0
Google
20.1
19.7
8.6
—
0.2
10.2
6.4
5.0
9.4
5.3
Meta
22.2
20.5
4.1
14.5
0.3
13.3
10.0
6.2
7.4
5.9
Reson8
23.7
24.4
14.3
16.7
0.7
10.7
5.8
3.7
7.7
6.0
Speechmatics
20.6
19.8
4.7
22.2
0.2
38.2
11.1
4.7
8.1
6.6
Soniox
18.2
21.5
4.6
31.8
0.6
11.3
7.9
7.6
8.2
6.7
AWS
19.3
28.1
10.2
27.5
1.3
9.1
8.9
3.9
8.7
7.2
ElevenLabs
18.5
20.4
7.0
—
0.2
13.0
9.3
5.1
17.1
7.4
Zoom
18.2
24.9
10.5
—
0.3
32.5
6.2
3.8
9.3
7.9
Mistral
23.1
20.8
12.3
—
0.0
18.5
9.5
6.6
18.9
10.0
Azure
29.7
33.1
14.4
20.4
0.2
14.6
12.7
5.5
9.2
10.2
Deepgram
26.1
25.9
13.2
34.0
0.2
17.1
12.7
6.5
9.6
10.4
xAI
24.0
26.1
15.7
25.7
1.2
20.6
14.8
5.8
9.4
11.6
OpenAI
38.1
20.2
17.2
—
0.1
33.3
12.3
8.0
15.9
12.5
Smallest
29.4
33.2
21.4
41.4
1.4
14.6
11.9
5.4
10.9
12.6
Gladia
30.6
24.9
17.0
—
4.9
17.1
16.2
7.0
10.7
14.1
Cartesia
40.1
27.8
21.1
—
3.1
50.5
18.7
7.7
10.6
16.1
BestWorstCells are coloured by rank within each column (solid blue is best, white is worst).
Figure 1Rank across 9 headline metrics, vendors ordered by mean rank.source = the tables below; each column’s provenance is given with its table
Position in the field
Voice-agent entity error14.41% · 1st of 22 systems · nearest competitor Soniox Realtime 18.19% −20.8%
Judge-scored entity error18.93% · 2nd of 22 systems · nearest competitor Google Gemini 3.5 Live 19.68% −3.8%
Medical entity error4.24% · 2nd of 22 systems · nearest competitor Meta Muse Voice Transcribe 4.08% +3.9%
Diarization DER16.68% · 3rd of 12 systems · nearest competitor Meta Muse Voice Transcribe 14.51% +15.0%
Non-speech response rate0.27% · 9th of 22 systems · nearest competitor Mistral Voxtral Mini 0.03% +0.2 pt
Code-switching (5 pairs)7.20% · 1st of 23 systems · nearest competitor AWS Transcribe 9.13% −21.1%
Multilingual (FLEURS)8.44% over 18 lang. · 5th of 22 systems · nearest competitor Reson8 5.77% over 7 lang. +46.3%
English short-form WER3.44% · 1st of 22 systems · nearest competitor Reson8 3.70% −7.0%
Long-form English WER8.82% · 6th of 22 systems · nearest competitor Meta Muse Voice Transcribe 7.37% +19.7%
We lead We trail The margin is blue where we lead, grey where we trail.
One row per headline metric from the grid above: our value, our rank among systems with a complete run, and the margin to the runner-up where we lead or to the leader where we do not. Margins are relative for error rates and absolute for rates below one percent.
02Voice agents
Conversational audio as a voice agent hears it, across 22 systems, three noise environments, 12 speaker groups and three utterance lengths. The headline metric, entity error (NEER), is how often the names, addresses and numbers an agent acts on come out wrong.
Table 1Entity error rate (%) by entity class. Colour is rank within the column. We lead in 9 of 11 columns.normalized entity error rate (NEER) · runs 2026-08-27 → 2026-09-22 · one row per model, most recent run · corpus value = mean of the three natural-noise environments
Entity error rate (%) by entity class. Colour is rank within the column. We lead in 9 of 11 columns.
System
Overall
Names
Places
Addresses
Phone
Email
Dates
Numbers
Codes/IDs
Medical
Technical
AssemblyAI Universal-3.6 Pro
14.4
10.9
6.0
29.0
2.4
17.2
18.6
12.9
10.0
16.1
20.9
AssemblyAI Universal-3.5 Pro
14.8
11.2
7.3
27.9
2.6
17.1
19.1
13.2
10.2
17.3
21.7
Soniox Realtime
18.2
16.6
7.7
27.9
2.7
22.7
30.8
19.7
8.2
19.1
26.5
Zoom Scribe (Live)
18.2
16.5
7.0
30.0
9.6
22.4
20.9
14.9
13.9
20.0
26.4
ElevenLabs Scribe v2
18.5
14.8
8.9
28.6
3.4
21.2
26.5
17.2
12.0
24.7
27.2
AWS Transcribe
19.3
21.3
7.3
26.8
3.1
21.9
27.8
18.8
13.0
19.0
33.7
Google Gemini 3.5 Live
20.1
15.8
9.3
32.4
7.3
23.9
30.4
17.3
16.9
20.5
27.4
Speechmatics Enhanced
20.6
23.8
6.0
31.6
3.2
24.2
31.1
17.1
19.4
17.4
31.8
Meta Muse Voice Transcribe
22.2
23.9
8.4
38.5
10.0
20.0
20.0
17.4
27.8
25.6
30.5
Mistral Voxtral Mini
23.1
18.3
10.3
38.9
7.7
27.5
28.6
17.3
16.1
31.6
34.3
Reson8
23.7
18.6
8.9
30.8
11.9
31.8
26.6
18.7
20.9
38.0
31.3
xAI Grok Streaming
24.0
17.9
11.8
40.2
12.4
30.0
19.9
18.1
20.0
35.7
34.3
Deepgram Nova-3
26.1
24.3
11.0
30.4
4.5
34.4
43.7
16.8
27.6
29.1
39.4
Smallest Pulse
29.4
27.5
15.4
41.8
12.1
39.5
24.9
19.5
33.2
38.5
42.0
Azure
29.7
32.4
17.9
53.6
8.7
22.2
30.5
23.5
35.4
32.4
40.5
Deepgram Flux EN
30.1
29.0
13.2
41.0
11.5
40.7
22.8
21.0
46.1
32.3
43.9
Deepgram Flux Multi
30.4
32.1
14.2
41.7
9.3
36.2
22.3
22.3
46.0
37.1
43.1
Gladia Solaria
30.6
21.3
13.5
39.6
20.6
42.4
31.6
19.2
37.4
40.5
40.0
Speechmatics Standard
30.6
30.6
9.3
41.6
6.8
57.1
32.5
20.1
43.3
28.8
36.2
OpenAI GPT-4o Transcribe
38.1
27.8
17.9
51.7
42.8
50.3
33.2
30.5
47.7
37.3
41.8
OpenAI GPT-4o mini Transcribe
38.9
29.3
19.6
53.3
40.6
52.1
35.6
28.3
46.0
42.5
41.9
Cartesia Ink-Whisper
40.1
25.1
14.6
56.9
58.9
47.1
33.6
25.3
52.6
45.1
42.1
BestWorstCells are coloured by rank within each column (solid blue is best, white is worst).
Background noise and utterance length
No background noise Natural low noise Natural heavy noise
Figure 2Entity error rate (%) by background-noise condition, sorted by the no-noise value. We lead in all 3 conditions.normalized entity error rate (NEER) · runs 2026-08-27 → 2026-09-22 · one row per model, most recent run
Listen to the test data
One caller turn per noise environment, quietest first, each with an entity an agent acts on: a name, an account number, an e-mail login, a medication. Words that differ from the human reference are marked.
1 of 8
Voice agentsno noise4 s · AssemblyAI voice-agent corpus — scripted caller scenarios voiced by real speakers
0:00 / 0:00
Four seconds, no noise. A short request where the other systems substitute unrelated words.
Reference
My cable modem keeps rebooting every 30-40 minutes, can you help?
AssemblyAI
My cable modem keeps rebooting every 30 to 40 minutes. Can you help?
ElevenLabs Scribe v2 Realtime
My key won't let them keep changing every 30 to 40 minutes. Can you help?
Soniox Realtime
Mark Kilmer and keeps removing every 30 to 40 minutes. Can you help?
Short utterances Medium Long
Figure 3Entity error rate (%) by utterance length. We lead in all 3 length bands.normalized entity error rate (NEER) · runs 2026-08-27 → 2026-09-22 · one row per model, most recent run
Speaker groups
Table 2Entity error rate (%) by speaker group. Colour is rank within the column. We lead in all 12 groups.normalized entity error rate (NEER) · runs 2026-08-27 → 2026-09-22 · one row per model, most recent run
Entity error rate (%) by speaker group. Colour is rank within the column. We lead in all 12 groups.
System
US Northeast
US South
US Gen. Am.
US AAVE
US Asian-Am.
US other
UK
Canada
Native other
Spanish-acc.
Indian-acc.
Non-nat. other
meanrank
AssemblyAI Universal-3.6 Pro
10.1
11.5
11.4
12.4
12.0
13.9
14.0
10.5
15.2
16.1
19.1
18.9
1.0
AssemblyAI Universal-3.5 Pro
10.1
11.6
11.6
12.8
12.2
14.5
14.2
11.0
15.2
16.6
20.3
19.1
2.0
Zoom Scribe (Live)
13.8
15.8
15.0
15.3
15.8
20.4
18.7
13.5
19.1
20.5
23.9
24.1
3.4
ElevenLabs Scribe v2
14.6
15.8
15.0
16.6
16.1
19.4
19.5
13.7
19.8
22.9
23.8
25.4
4.2
Soniox Realtime
16.3
16.5
16.5
18.1
18.9
20.7
19.3
15.8
20.5
22.4
23.8
24.8
5.3
Google Gemini 3.5 Live
17.2
18.6
17.8
20.7
16.8
22.4
21.5
15.6
20.6
23.6
21.2
25.8
6.6
AWS Transcribe
17.7
18.4
19.2
20.7
19.8
22.4
20.1
18.0
21.6
23.8
24.4
27.4
7.7
xAI Grok Streaming
15.0
19.4
18.3
20.6
17.7
23.2
23.0
15.4
24.4
22.4
27.6
28.2
8.0
Mistral Voxtral Mini
18.0
19.0
18.3
20.4
19.6
22.4
20.6
15.5
23.6
28.7
28.0
31.1
8.8
Reson8
17.7
20.5
19.2
19.7
20.0
25.6
21.6
16.7
23.4
27.7
28.5
32.0
9.8
Meta Muse Voice Transcribe
19.5
20.3
20.6
21.5
21.1
22.7
22.7
19.2
24.1
28.0
26.6
29.6
10.5
Speechmatics Enhanced
21.4
21.0
21.2
22.0
22.1
24.8
22.6
19.8
23.8
27.0
26.6
29.6
10.8
Gladia Solaria
22.0
25.2
24.6
27.3
27.3
29.6
28.0
20.5
30.3
33.5
32.7
37.6
14.0
Deepgram Nova-3
24.4
26.4
26.2
27.0
27.2
28.4
28.6
24.3
30.1
31.0
32.9
35.1
14.3
Smallest Pulse
21.6
25.5
24.7
26.0
29.4
30.1
26.5
21.5
30.2
34.6
31.0
40.1
14.7
Deepgram Flux EN
24.4
26.5
26.0
27.5
27.1
30.4
32.2
25.7
32.3
33.7
34.9
37.6
15.8
Azure
26.6
29.0
28.4
29.0
28.9
32.1
31.4
26.0
31.9
36.4
35.8
37.8
17.5
Deepgram Flux Multi
25.3
27.2
27.1
29.2
27.7
32.6
33.0
26.3
33.3
36.1
36.3
39.8
17.8
Speechmatics Standard
26.7
28.9
28.4
31.6
32.0
33.4
31.5
28.2
33.5
37.4
35.7
39.7
18.8
OpenAI GPT-4o Transcribe
28.9
32.7
30.0
31.4
30.7
35.7
34.9
25.1
36.3
42.3
39.8
44.2
19.9
Cartesia Ink-Whisper
27.9
32.2
30.5
31.9
32.7
35.7
34.0
25.8
36.8
41.6
40.9
44.1
20.4
OpenAI GPT-4o mini Transcribe
29.4
33.0
30.5
32.3
32.4
37.0
35.3
27.5
37.5
42.9
41.6
46.9
21.8
BestWorstCells are coloured by rank within each column (solid blue is best, white is worst).
Ablations on our own model
wordboost OFF wordboost ON
0%10%20%30%0%
11.25%
Overall
4.36%
Names
6.14%
Places
28.16%
Addresses
2.69%
Phone
15.29%
Email
20.01%
Dates
12.91%
Numbers
9.85%
Codes/IDs
7.16%
Medical
5.90%
Technical
5.05%
WER
Figure 4Word boost (custom vocabulary) on versus off, same model and corpus. Change in entity error with word boost on: names −60.9%, technical terms −72.8%.normalized entity error rate (NEER) + WER, wordboost OFF vs ON · runs 2026-09-22 · one row per model, most recent run
agent context OFF agent context ON
0%10%20%30%0%
15.35%
Overall
10.43%
Names
5.58%
Places
29.70%
Addresses
2.44%
Phone
16.30%
Email
20.95%
Dates
15.61%
Numbers
10.45%
Codes/IDs
17.57%
Medical
24.40%
Technical
5.61%
WER
Figure 5Agent context on versus off: the agent’s previous turn is passed to the recognizer.normalized entity error rate (NEER) + WER, same model with and without agent context · runs 2026-08-27 → 2026-08-28 · one row per model, most recent run
Word error rate
Table 3WER by noise condition.normalized word error rate · runs 2026-08-27 → 2026-09-22 · one row per model, most recent run
WER by noise condition.
#
System model / snapshot
Mean 3 conditions
No noise
Low noise
Heavy noise
1
AssemblyAI Universal-3.6 Pro Realtime
5.19%
4.37%
5.08%
6.13%
2
AssemblyAI Universal-3.5 Pro Realtime
5.72%
4.61%
5.47%
7.08%
3
Soniox Realtime
6.83%
5.98%
6.84%
7.68%
4
Meta Muse Voice Transcribe
7.48%
6.92%
7.50%
8.02%
5
ElevenLabs Scribe v2 Realtime
7.78%
6.18%
7.53%
9.64%
6
Speechmatics Enhanced Realtime
8.14%
6.96%
7.96%
9.49%
7
Google Gemini 3.5 Transcribe Live
8.48%
7.44%
8.49%
9.52%
8
AWS Transcribe Streaming
8.58%
7.27%
8.27%
10.19%
9
Deepgram Nova-3
8.64%
7.53%
8.45%
9.93%
10
Zoom Scribe (Live)
9.36%
8.04%
9.29%
10.74%
11
Reson8 Reson8
9.41%
8.21%
9.38%
10.64%
12
Mistral Voxtral Mini Realtime
10.43%
9.25%
10.45%
11.59%
13
Smallest Pulse
10.74%
9.85%
10.94%
11.43%
14
Azure Realtime STT
10.86%
9.78%
10.65%
12.15%
15
Speechmatics Standard Realtime
12.05%
10.47%
11.77%
13.91%
16
Deepgram Flux EN
13.50%
12.29%
13.36%
14.85%
17
Deepgram Flux Multi
14.26%
12.85%
14.06%
15.87%
18
Cartesia Ink-Whisper
15.17%
13.71%
15.00%
16.81%
19
Gladia Solaria
16.25%
12.99%
16.04%
19.71%
20
OpenAI GPT-4o Transcribe
18.37%
17.05%
18.44%
19.63%
21
OpenAI GPT-4o mini Transcribe
19.55%
18.12%
19.77%
20.76%
22
xAI Grok Streaming
20.28%
18.45%
19.89%
22.50%
Best Others The best value in each column is filled blue.
Data
●Caller corpus 12,460 clips
All entity classes pooled into one number: how often an agent-critical value (a name, address, number or identifier) comes back wrong.
●Noise environments
Three levels of real noise, recorded rather than mixed in. None: room and handset noise only. Low: incidental noise from a home, office or parked car. Heavy: noise that competes with the voice.
●Utterance length
Short turns (confirmations, single values, one-word answers), medium and long. Short turns give the model almost no context.
●Speaker groups
Fourteen groups: US Northeast, South, General American, AAVE, Asian-American and other; UK; Canada; native speakers overall; Spanish- and Indian-accented and other non-native speakers.
●Entity classes
Names (including letter-by-letter spellings), places, street addresses, phone numbers, email addresses, dates and times, numbers and currency, codes and identifiers, medical and technical terms.
○public, open source ◆public, licensed ●private, ours
03Healthcare
Medical entity error is measured on primary-care consultations (PriMock57), licensed clinical dictation and respiratory-clinic recordings.
Overall
Table 4Medical entity error rate. We lead on primary care.normalized medical entity error rate (MER) · runs 2026-08-18 → 2026-09-22 · one row per model, most recent run
Medical entity error rate. We lead on primary care.
#
System model / snapshot
Mean 3 sets
Primary care PriMock57
Dictation
Respiratory
1
Meta Muse Voice Transcribe
4.08%
7.18%
3.09%
1.97%
2
AssemblyAI Universal-3.6 Pro Realtime
4.24%
6.23%
4.20%
2.28%
3
Soniox Realtime
4.58%
7.18%
4.04%
2.51%
4
Speechmatics Enhanced Realtime
4.66%
8.13%
3.65%
2.21%
5
AssemblyAI Universal-3.5 Pro Realtime
4.81%
7.07%
5.08%
2.28%
6
ElevenLabs Scribe v2 Realtime
7.00%
10.89%
8.01%
2.09%
7
Google Gemini 3.5 Transcribe Live
8.57%
15.45%
7.69%
2.56%
8
AWS Transcribe Streaming
10.21%
11.09%
16.49%
3.06%
9
Zoom Scribe (Live)
10.50%
15.95%
8.09%
7.47%
10
Mistral Voxtral Mini Realtime
12.26%
12.67%
16.02%
8.09%
11
Deepgram Nova-3
13.22%
19.01%
17.37%
3.29%
12
Speechmatics Standard Realtime
14.24%
18.68%
18.72%
5.32%
13
Deepgram Flux EN
14.28%
20.49%
18.32%
4.02%
14
Reson8 Reson8
14.31%
16.37%
21.81%
4.76%
15
Azure Realtime STT
14.41%
17.21%
19.11%
6.92%
16
xAI Grok Streaming
15.69%
17.95%
19.35%
9.78%
17
Gladia Solaria
17.03%
18.83%
25.30%
6.97%
18
OpenAI GPT Realtime
17.21%
13.73%
5.39%
32.50%
19
Cartesia Ink-Whisper
21.09%
24.39%
31.09%
7.78%
20
Smallest Pulse
21.40%
25.42%
30.77%
8.01%
21
OpenAI GPT-4o Transcribe
26.56%
23.34%
14.91%
41.43%
22
OpenAI GPT-4o mini Transcribe
27.94%
25.45%
19.27%
39.11%
Best Others The best value in each column is filled blue.
Medication and disease
Table 5Medication and disease terms scored separately.medication / disease entity error rate · runs 2026-08-18 → 2026-09-22 · one row per model, most recent run
Medication and disease terms scored separately.
#
System model / snapshot
Medication mean
Primary care
Dictation
Respiratory
Disease mean
Primary care
Dictation
Respiratory
1
Speechmatics Enhanced Realtime
6.62%
8.92%
5.04%
5.90%
3.38%
7.35%
0.93%
1.86%
2
Meta Muse Voice Transcribe
6.66%
10.40%
4.20%
5.38%
2.04%
3.99%
0.93%
1.19%
3
AssemblyAI Universal-3.6 Pro Realtime
7.11%
9.13%
5.52%
6.67%
2.11%
3.36%
1.64%
1.34%
4
Soniox Realtime
7.20%
10.83%
3.84%
6.92%
3.22%
3.57%
4.44%
1.64%
5
AssemblyAI Universal-3.5 Pro Realtime
7.37%
9.34%
6.60%
6.15%
2.78%
4.83%
2.10%
1.42%
6
ElevenLabs Scribe v2 Realtime
10.11%
13.57%
11.64%
5.13%
3.53%
8.17%
0.93%
1.49%
7
Google Gemini 3.5 Transcribe Live
12.23%
19.19%
10.80%
6.70%
5.16%
11.76%
1.64%
2.09%
8
AWS Transcribe Streaming
14.40%
15.07%
21.97%
6.15%
4.90%
7.14%
5.84%
1.72%
9
Zoom Scribe (Live)
15.21%
19.11%
11.40%
15.13%
6.31%
12.82%
1.64%
4.47%
10
Deepgram Nova-3
17.76%
25.05%
20.53%
7.69%
8.83%
13.03%
11.21%
2.24%
11
OpenAI GPT Realtime
17.76%
17.83%
6.72%
28.72%
14.02%
9.66%
2.80%
29.60%
12
Mistral Voxtral Mini Realtime
18.22%
18.47%
22.09%
14.10%
5.97%
6.93%
4.21%
6.79%
13
Azure Realtime STT
18.60%
23.57%
21.97%
10.26%
9.57%
10.92%
13.55%
4.25%
14
xAI Grok Streaming
19.50%
24.63%
26.05%
7.81%
7.74%
11.34%
6.31%
5.56%
15
Deepgram Flux EN
19.77%
29.30%
20.77%
9.23%
9.18%
11.76%
13.55%
2.24%
16
Speechmatics Standard Realtime
19.81%
26.98%
21.13%
11.31%
9.56%
10.53%
14.02%
4.13%
17
Reson8 Reson8
21.23%
23.78%
27.85%
12.05%
7.23%
9.00%
10.05%
2.61%
18
Gladia Solaria
23.23%
23.81%
32.41%
13.46%
9.69%
12.86%
11.45%
4.75%
19
Smallest Pulse
26.83%
31.55%
41.42%
7.53%
12.04%
18.30%
10.05%
7.76%
20
Cartesia Ink-Whisper
28.31%
33.55%
37.01%
14.36%
13.72%
15.34%
19.57%
6.26%
21
OpenAI GPT-4o Transcribe
28.38%
27.18%
16.93%
41.03%
22.68%
19.54%
10.98%
37.51%
22
OpenAI GPT-4o mini Transcribe
31.72%
32.91%
21.49%
40.77%
22.84%
18.07%
14.95%
35.50%
Best Others The best value in each column is filled blue.
Public corpus of mock primary-care consultations: clinicians and actor-patients work through realistic appointments, recorded as telemedicine calls.
◆Medical dictation 93 recordings
Licensed clinician dictation, notes and findings spoken for the record, dense with medication and disease terms.
●Respiratory clinic recordings 272 recordings
Real clinical conversation from respiratory clinic encounters, with the accents, interruptions and room noise of a working clinic.
○public, open source ◆public, licensed ●private, ours
04English
General English: entity accuracy and non-speech handling first, then word error rate on short-form, accented and long-form sets.
Table 6Judge-scored entity error (%) on general English, by class. We lead on both overall scores. Per class, we lead on person names, organisations and URLs. A 0.00 cell means no data.LLM-judge-scored entity error rate (normalized and formatted) · runs 2026-08-18 → 2026-09-22 · one row per model, most recent run
Judge-scored entity error (%) on general English, by class. We lead on both overall scores. Per class, we lead on person names, organisations and URLs. A 0.00 cell means no data.
System
All classes
All, formatted
Persons
Products
Addresses
Orgs
Places
Phone
Email
URLs
Dates
Credit cards
Alphanum.
AssemblyAI Universal-3.5 Pro
18.1
27.4
19.9
19.2
21.5
16.4
20.5
9.3
16.7
15.1
16.5
19.6
9.0
AssemblyAI Universal-3.6 Pro
18.9
27.6
20.3
22.6
24.3
17.1
21.2
12.7
16.7
14.9
20.7
11.3
7.7
Google Gemini 3.5 Live
19.7
29.5
22.9
22.1
30.6
17.6
18.1
8.9
19.0
15.4
15.6
22.0
12.3
Speechmatics Enhanced
19.8
37.5
25.2
22.2
24.6
21.7
16.5
0.4
25.8
21.5
15.8
n/d
7.7
OpenAI GPT Realtime
20.2
38.6
24.8
27.5
21.4
20.0
17.7
2.1
19.9
21.9
19.0
3.0
13.4
ElevenLabs Scribe v2
20.4
29.2
24.8
22.6
29.2
22.6
18.9
n/d
17.2
14.9
19.4
n/d
8.7
Meta Muse Voice Transcribe
20.5
44.4
24.4
20.7
38.2
18.3
17.5
2.5
15.8
29.4
12.7
2.4
10.3
Mistral Voxtral Mini
20.8
31.7
27.8
24.0
26.4
19.3
17.4
3.8
13.1
30.0
16.4
n/d
7.5
Soniox Realtime
21.5
33.6
31.6
31.9
20.1
21.9
17.8
1.7
11.3
19.7
15.7
—
8.8
OpenAI GPT-4o Transcribe
24.4
37.8
26.2
25.9
33.0
23.3
26.2
18.1
25.3
19.7
21.9
21.4
10.5
Reson8
24.4
37.1
37.3
24.2
25.9
23.6
18.5
8.0
16.3
26.9
19.2
4.2
8.3
Gladia Solaria
24.9
41.2
32.5
26.6
29.2
24.1
22.1
4.2
21.3
31.0
20.2
7.7
10.3
Zoom Scribe (Live)
24.9
39.3
34.3
29.6
25.9
23.9
25.1
—
15.4
25.5
23.6
—
5.7
OpenAI GPT-4o mini Transcribe
25.0
38.1
28.7
27.1
33.1
24.0
27.2
8.9
24.0
20.1
20.2
8.3
12.6
Deepgram Nova-3
25.9
42.4
36.1
28.0
24.8
25.6
22.3
16.9
16.3
20.8
29.3
7.1
8.8
xAI Grok Streaming
26.1
50.2
34.0
28.8
27.6
24.8
26.1
2.5
19.9
28.3
24.2
4.8
8.5
Speechmatics Standard
26.6
44.4
33.9
25.5
29.1
26.2
20.4
6.3
30.8
39.8
16.5
n/d
26.5
Cartesia Ink-Whisper
27.8
43.0
37.6
28.2
32.9
27.3
26.1
12.7
18.9
24.2
20.9
13.7
10.7
AWS Transcribe
28.1
43.8
46.2
29.6
20.4
27.6
23.4
0.4
20.8
28.3
17.3
8.3
7.5
Deepgram Flux EN
31.4
62.5
49.1
33.0
32.7
30.2
25.0
3.4
15.8
31.4
23.5
n/d
10.0
Azure
33.1
97.5
45.0
35.9
50.2
34.4
27.8
2.1
14.5
24.7
31.0
0.6
10.0
Smallest Pulse
33.2
58.0
53.0
33.0
33.0
31.7
28.7
3.8
17.6
25.3
21.1
5.4
10.0
BestWorstCells are coloured by rank within each column (solid blue is best, white is worst).
Figure 6Non-speech response rate: how often a system produces text for a clip with no speech. Lower is better. The axis is clipped at 5%; the notched bar exceeds it.non-speech response rate (share of pure non-speech clips that produce any text) · runs 2026-08-18 → 2026-09-22 · one row per model, most recent run
Word error rate — short-form
Table 7Short-form English WER. We lead on the mean and Common Voice. Systems without a Common Voice run have no mean.normalized word error rate · runs 2026-08-18 → 2026-09-22 · one row per model, most recent run
Short-form English WER. We lead on the mean and Common Voice. Systems without a Common Voice run have no mean.
#
System model / snapshot
Mean 3 sets
Common Voice
LibriSpeech clean
LibriSpeech other
1
AssemblyAI Universal-3.6 Pro Realtime
3.44%
5.18%
1.88%
3.25%
2
Reson8 Reson8
3.70%
6.98%
1.43%
2.68%
3
Zoom Scribe (Live)
3.75%
6.32%
1.72%
3.21%
4
AssemblyAI Universal-3.5 Pro Realtime
3.83%
6.44%
1.83%
3.22%
5
AWS Transcribe Streaming
3.93%
6.12%
1.90%
3.77%
6
Speechmatics Enhanced Realtime
4.69%
6.78%
2.31%
4.99%
7
Google Gemini 3.5 Transcribe Live
5.04%
8.22%
2.07%
4.83%
8
ElevenLabs Scribe v2 Realtime
5.13%
9.01%
1.96%
4.43%
9
Smallest Pulse
5.41%
9.34%
2.19%
4.69%
10
Azure Realtime STT
5.54%
8.68%
2.44%
5.49%
11
xAI Grok Streaming
5.81%
11.04%
2.04%
4.36%
12
Meta Muse Voice Transcribe
6.20%
11.24%
2.15%
5.22%
13
Speechmatics Standard Realtime
6.41%
9.43%
3.18%
6.61%
14
Deepgram Flux EN
6.53%
7.68%
3.56%
8.35%
15
Mistral Voxtral Mini Realtime
6.61%
12.23%
2.10%
5.49%
16
Gladia Solaria
7.01%
12.88%
2.73%
5.42%
17
Deepgram Nova-3
7.46%
12.38%
3.28%
6.72%
18
Soniox Realtime
7.60%
12.38%
3.27%
7.15%
19
Cartesia Ink-Whisper
7.66%
13.31%
3.19%
6.48%
20
OpenAI GPT Realtime
7.95%
13.83%
2.59%
7.44%
21
OpenAI GPT-4o Transcribe
8.11%
14.99%
2.20%
7.15%
22
OpenAI GPT-4o mini Transcribe
8.75%
15.76%
2.53%
7.97%
Best Others The best value in each column is filled blue.
Listen to the test data
Two read sentences from the short-form sets above, human reference beside each transcript. Words that differ from it are marked.
1 of 2
Short-form English17 s · Common Voice (CC0)
0:00 / 0:00
A read sentence. The other system appends a sentence and a name that are not in the audio.
Reference
You can do a lot just using your voice, but there are still a few times you'll find yourself reaching for a mouse.
AssemblyAI
You can do a lot just using your voice, but there are still a few times you'll find yourself reaching for a mouse.
Speechmatics Enhanced
I know you can do a lot just using your voice, but there are still a few times you'll find yourself reaching for a mouse. And I'm Larry Hanson. We're here in this beautiful hall.
Word error rate — accented English
Table 8Accented English WER. We lead on British-accented English and Indian-accented English.normalized word error rate · runs 2026-08-18 → 2026-09-22 · one row per model, most recent run
Accented English WER. We lead on British-accented English and Indian-accented English.
#
System model / snapshot
Mean 4 sets
British
French-acc.
Spanish-acc.
Indian-acc.
1
Azure Realtime STT
6.09%
6.25%
4.44%
5.52%
8.15%
2
AssemblyAI Universal-3.6 Pro Realtime
7.29%
5.53%
8.54%
9.62%
5.47%
3
AssemblyAI Universal-3.5 Pro Realtime
7.46%
5.72%
8.69%
9.53%
5.89%
4
Zoom Scribe (Live)
7.94%
6.14%
9.57%
10.03%
6.04%
5
Reson8 Reson8
9.24%
5.76%
11.89%
12.60%
6.73%
6
Google Gemini 3.5 Transcribe Live
9.39%
6.30%
11.86%
12.82%
6.59%
7
Meta Muse Voice Transcribe
9.64%
7.16%
11.26%
12.22%
7.93%
8
xAI Grok Streaming
9.85%
7.33%
12.21%
12.88%
6.99%
9
ElevenLabs Scribe v2 Realtime
10.10%
6.92%
12.12%
13.29%
8.07%
10
OpenAI GPT-4o Transcribe
10.30%
7.27%
11.83%
13.89%
8.22%
11
Speechmatics Enhanced Realtime
10.44%
6.70%
13.26%
14.80%
6.98%
12
Gladia Solaria
11.07%
8.67%
12.60%
13.85%
9.15%
13
AWS Transcribe Streaming
11.10%
7.02%
15.49%
15.47%
6.42%
14
Cartesia Ink-Whisper
11.23%
8.16%
12.92%
14.57%
9.27%
15
OpenAI GPT-4o mini Transcribe
11.36%
7.88%
13.31%
15.35%
8.88%
16
Soniox Realtime
11.69%
8.99%
12.87%
14.32%
10.58%
17
OpenAI GPT Realtime
11.78%
8.31%
14.83%
15.78%
8.20%
18
Smallest Pulse
11.85%
7.70%
15.15%
16.53%
8.03%
19
Deepgram Nova-3
11.95%
9.14%
13.92%
15.19%
9.57%
20
Deepgram Flux EN
12.55%
9.61%
14.83%
16.12%
9.64%
21
Speechmatics Standard Realtime
12.58%
8.43%
16.29%
17.46%
8.14%
22
Mistral Voxtral Mini Realtime
12.63%
7.39%
14.19%
15.87%
13.06%
Best Others The best value in each column is filled blue.
Word error rate — long-form
Table 9Long-form English WER. TED-LIUM references omit sponsor read-outs present in the audio.normalized word error rate · runs 2026-08-18 → 2026-09-22 · one row per model, most recent run
Long-form English WER. TED-LIUM references omit sponsor read-outs present in the audio.
#
System model / snapshot
Mean 3 sets
TED-LIUM
Rev16
CALLHOME per channel
1
Meta Muse Voice Transcribe
7.37%
6.01%
6.57%
8.36%
2
Reson8 Reson8
7.74%
6.01%
6.87%
9.17%
3
Speechmatics Enhanced Realtime
8.08%
6.29%
7.34%
9.59%
4
Soniox Realtime
8.19%
7.13%
8.17%
9.46%
5
AWS Transcribe Streaming
8.71%
6.45%
7.81%
9.82%
6
AssemblyAI Universal-3.6 Pro Realtime
8.82%
8.14%
7.97%
11.07%
7
Speechmatics Standard Realtime
8.90%
6.83%
8.06%
10.77%
8
Azure Realtime STT
9.20%
7.16%
8.17%
10.87%
9
Zoom Scribe (Live)
9.34%
7.41%
7.17%
13.26%
10
AssemblyAI Universal-3.5 Pro Realtime
9.35%
8.58%
8.24%
11.69%
11
Google Gemini 3.5 Transcribe Live
9.42%
6.87%
8.04%
10.99%
12
xAI Grok Streaming
9.43%
6.30%
8.00%
12.82%
13
Deepgram Flux EN
9.59%
7.01%
8.08%
12.12%
14
Deepgram Nova-3
9.75%
7.43%
9.40%
11.19%
15
Cartesia Ink-Whisper
10.60%
7.13%
10.42%
15.62%
16
Gladia Solaria
10.72%
7.75%
9.04%
18.25%
17
Smallest Pulse
10.95%
5.23%
7.62%
14.13%
18
OpenAI GPT Realtime
15.87%
5.70%
6.02%
20.90%
19
ElevenLabs Scribe v2 Realtime
17.12%
5.96%
8.38%
45.51%
20
Mistral Voxtral Mini Realtime
18.93%
6.96%
8.33%
52.91%
21
OpenAI GPT-4o mini Transcribe
19.10%
7.28%
21.26%
18.56%
22
OpenAI GPT-4o Transcribe
27.45%
9.88%
28.51%
30.33%
Best Others The best value in each column is filled blue.
Data
●Judge-scored entities
The short-form English audio below. An LLM judge grades each entity (person, product, organisation, place, address, phone, email, date, ID, URL) on whether the transcript names the same real-world thing, with and without formatting.
●Non-speech 2,950 clips
Hold music, line noise and silence, with no words at all. The score is how often a model emits text anyway.
○Common Voice 16,029 clips
Crowdsourced read sentences from Mozilla’s public Common Voice, recorded by volunteers on their own phones and laptops, so accents, microphones and rooms vary widely.
○LibriSpeech test-clean 2,620 utterances
Read English audiobook speech from the public LibriSpeech corpus, recorded close-mic by LibriVox volunteers. The "clean" split is the easier half: clear speakers, little noise.
○LibriSpeech test-other 2,939 utterances
The harder half of the same LibriSpeech set: speakers and recordings that automatic transcription finds more difficult, still read prose.
◆Accented English 9,315 files
Four licensed sets of prompted British-, French-, Spanish- and Indian-accented English: accuracy outside the General American accent most models are tuned on.
○TED-LIUM 11 full-length talks
Full TED talks from the public TED-LIUM corpus: one prepared speaker on stage, recorded end to end, many minutes of continuous speech per file.
○Rev16 9 full-length recordings
Podcast-style recordings from the public Rev16 set: spontaneous multi-speaker studio conversation, episode length.
◆CALLHOME 71 scored channels
Public CALLHOME telephone calls between friends and family: unscripted, fast, overlapping conversation over a real phone line. Each speaker's channel is scored on its own.
○public, open source ◆public, licensed ●private, ours
05Multilingual
Eighteen languages on FLEURS and Common Voice, drawn from the thirty-two Universal-3.6 Pro Realtime supports. A system is scored on every language it supports, and a cross marks one its vendor does not document for realtime use — so the means below cover different numbers of languages and the Languages column says how many.
FLEURS — Systems covering all 18 languages
Table 10WER on FLEURS (CER for languages marked *). Each mean covers only the languages that system supports; the Languages column gives the count.normalized word error rate (character error rate for Japanese, Chinese, Korean) · runs 2026-08-25 → 2026-09-23 · one row per model, most recent run
WER on FLEURS (CER for languages marked *). Each mean covers only the languages that system supports; the Languages column gives the count.
System
Mean over supported langs
Languages of 18
Spanish
French
German
Italian
Portuguese
Dutch
Danish
Swedish
Finnish
Russian
Turkish
Arabic
Hebrew
Hindi
Vietnamese
Japanese*
Chinese*
Korean*
Google Gemini 3.5 Live
6.4
18
3.8
5.4
4.7
3.2
4.7
5.0
9.1
7.3
7.5
5.3
5.3
6.7
18.9
5.0
5.2
5.0
9.4
3.8
Soniox Realtime
7.9
18
5.0
7.8
6.9
4.9
6.2
7.7
11.4
10.3
8.0
7.3
7.3
7.2
19.2
6.3
7.0
5.8
10.3
4.4
AssemblyAI Universal-3.6 Pro
8.4
18
4.1
5.0
5.0
4.0
5.5
5.9
12.7
6.6
10.8
7.5
8.6
12.1
18.2
7.5
13.9
6.3
14.0
4.3
AWS Transcribe
8.9
18
5.3
7.6
8.0
4.6
8.2
7.1
11.5
10.0
9.8
9.4
8.6
8.5
22.6
6.3
9.3
7.0
12.3
4.9
ElevenLabs Scribe v2
9.3
18
4.6
6.9
5.8
4.3
6.4
7.0
12.5
10.9
9.5
6.6
7.7
10.8
27.0
15.8
8.8
8.6
10.1
4.6
Meta Muse Voice Transcribe
10.0
18
6.5
8.0
6.4
6.0
8.7
8.6
13.6
11.0
10.7
10.1
10.5
8.9
27.0
8.1
6.5
9.8
13.8
6.1
Speechmatics Enhanced
11.1
18
4.7
6.5
6.8
5.4
7.3
8.5
11.8
9.9
8.0
10.0
9.1
6.5
18.5
7.0
8.6
31.3
33.8
5.4
Speechmatics Standard
11.6
18
5.4
8.1
8.2
5.7
7.3
9.5
11.7
10.2
8.7
10.5
8.7
8.4
19.3
7.0
9.3
31.4
34.7
5.2
OpenAI GPT Realtime
12.3
18
5.7
11.8
11.2
5.9
8.8
10.1
14.7
12.4
12.4
10.4
14.3
10.5
34.5
10.4
11.4
13.4
16.3
7.1
Azure
12.7
18
7.7
11.4
9.1
8.3
10.9
12.9
19.8
16.5
15.6
13.1
15.2
12.5
27.6
8.6
9.4
8.6
11.5
9.4
OpenAI GPT-4o Transcribe
12.8
18
4.3
10.6
13.1
5.5
8.5
13.0
17.7
10.1
9.2
8.4
14.8
9.4
34.7
12.4
18.7
16.1
17.4
6.7
OpenAI GPT-4o mini Transcribe
15.6
18
5.5
13.2
14.9
6.9
10.9
16.8
20.6
12.7
12.4
9.9
17.5
11.7
45.0
16.8
23.1
15.4
19.3
7.6
Gladia Solaria
16.2
18
6.7
12.7
10.9
7.5
8.7
11.6
22.6
15.3
15.9
10.3
12.2
15.1
45.4
23.9
21.5
23.6
21.1
6.6
Cartesia Ink-Whisper
18.7
18
9.0
16.9
12.7
9.7
10.7
17.4
25.4
16.1
17.0
15.7
16.2
18.1
42.0
27.5
23.6
20.4
26.3
10.9
BestWorstCells are coloured by rank within each column (solid blue is best, white is worst).
FLEURS — Systems covering fewer than 18
Table 11The same measurement on the systems that do not cover all eighteen. Each mean is over that system’s own languages, so these are not comparable with the table above or with each other — the Languages column gives the count, and a cross marks a language the vendor does not document for realtime use.normalized word error rate (character error rate for Japanese, Chinese, Korean) · runs 2026-08-25 → 2026-09-23 · one row per model, most recent run
The same measurement on the systems that do not cover all eighteen. Each mean is over that system’s own languages, so these are not comparable with the table above or with each other — the Languages column gives the count, and a cross marks a language the vendor does not document for realtime use.
BestWorstCells are coloured by rank within each column (solid blue is best, white is worst).
Common Voice — Systems covering all 18 languages
Table 12WER on Common Voice (CER for languages marked *). Each mean covers only the languages that system supports; the Languages column gives the count.normalized word error rate (character error rate for Japanese, Chinese, Korean) · runs 2026-08-25 → 2026-09-23 · one row per model, most recent run
WER on Common Voice (CER for languages marked *). Each mean covers only the languages that system supports; the Languages column gives the count.
System
Mean over supported langs
Languages of 18
Spanish
French
German
Italian
Portuguese
Dutch
Danish
Swedish
Finnish
Russian
Turkish
Arabic
Hebrew
Hindi
Vietnamese
Japanese*
Chinese*
Korean*
Speechmatics Enhanced
7.3
18
2.5
7.2
3.7
4.7
5.2
2.4
7.3
4.0
3.4
7.1
7.3
6.1
9.4
3.9
2.8
23.2
25.1
7.2
AWS Transcribe
7.5
18
4.1
8.8
5.8
4.2
4.9
2.7
7.3
5.9
5.7
4.8
10.3
9.6
15.0
5.1
8.2
15.9
8.5
7.5
AssemblyAI Universal-3.6 Pro
7.6
18
3.2
6.9
3.6
3.3
4.9
2.6
9.4
4.8
9.3
8.1
11.0
9.4
12.3
5.3
10.5
16.2
9.2
6.8
Speechmatics Standard
9.0
18
4.2
10.2
6.2
5.3
6.4
3.0
8.1
5.7
5.9
7.7
6.5
11.1
11.8
4.0
6.2
25.4
25.8
8.3
Google Gemini 3.5 Live
9.2
18
4.0
9.3
5.2
4.2
5.8
3.3
11.5
9.7
11.5
6.5
10.1
9.6
17.8
4.7
9.3
21.7
13.3
7.5
Soniox Realtime
9.5
18
6.0
12.2
7.6
6.3
7.4
4.6
9.5
8.6
7.7
9.0
10.8
10.2
12.5
6.6
10.1
19.2
14.1
8.7
Azure
10.0
18
6.8
11.0
6.1
6.2
7.2
4.6
11.9
10.0
11.1
7.0
12.3
11.5
20.0
4.9
8.8
19.8
9.9
10.4
Meta Muse Voice Transcribe
10.2
18
4.9
10.4
6.6
5.5
6.4
5.7
14.6
11.1
13.6
7.2
9.6
9.3
18.4
5.8
9.0
20.3
16.4
9.6
OpenAI GPT Realtime
12.3
18
6.9
12.6
9.3
7.8
9.4
6.8
14.5
12.0
11.8
8.2
11.1
12.7
28.9
8.4
13.4
20.2
16.4
10.5
ElevenLabs Scribe v2
14.6
18
5.8
10.2
7.1
6.1
9.7
5.5
14.5
12.9
14.6
10.2
15.0
20.6
43.4
21.6
21.5
20.7
12.1
11.0
OpenAI GPT-4o Transcribe
15.0
18
8.9
15.5
10.3
9.2
13.2
7.8
18.0
12.6
11.2
9.3
15.1
17.2
34.9
10.0
20.8
22.0
21.4
12.1
Gladia Solaria
16.1
18
7.8
14.4
10.5
9.8
9.5
5.7
17.5
12.6
14.0
9.5
16.2
19.8
36.1
22.0
20.6
26.2
29.0
9.6
OpenAI GPT-4o mini Transcribe
17.7
18
9.8
16.8
11.7
10.6
14.3
11.9
21.9
17.1
13.3
11.9
14.3
22.5
42.9
21.5
23.9
22.3
19.3
12.2
Cartesia Ink-Whisper
18.9
18
10.4
17.7
12.4
11.1
11.5
9.7
21.1
16.3
17.7
12.3
18.0
22.0
39.1
27.9
22.2
28.3
29.2
13.8
BestWorstCells are coloured by rank within each column (solid blue is best, white is worst).
Common Voice — Systems covering fewer than 18
Table 13The same measurement on the systems that do not cover all eighteen. Each mean is over that system’s own languages, so these are not comparable with the table above or with each other — the Languages column gives the count, and a cross marks a language the vendor does not document for realtime use.normalized word error rate (character error rate for Japanese, Chinese, Korean) · runs 2026-08-25 → 2026-09-23 · one row per model, most recent run
The same measurement on the systems that do not cover all eighteen. Each mean is over that system’s own languages, so these are not comparable with the table above or with each other — the Languages column gives the count, and a cross marks a language the vendor does not document for realtime use.
BestWorstCells are coloured by rank within each column (solid blue is best, white is worst).
Data
○FLEURS
Public corpus of parallel sentences read by native speakers in quiet settings, one test set per language. Shown: Spanish, French, German, Italian, Portuguese, Dutch, Hindi, Japanese and Chinese.
○Common Voice
Mozilla’s public Common Voice, per language: volunteers reading sentences on their own devices, so accents and microphones vary. Japanese and Chinese are scored by character error rate.
○public, open source ◆public, licensed ●private, ours
06Code-switching
Speech that switches between English and a second language within one utterance: five language pairs plus a Spanish–English conversational corpus.
Table 14Code-switching WER. We lead on the five-pair mean. We lead in all 5 pairs.normalized word error rate · runs 2026-08-18 → 2026-09-23 · one row per model, most recent run · pair value = mean of 4 sub-sets
Code-switching WER. We lead on the five-pair mean. We lead in all 5 pairs.
#
System model / snapshot
Mean 5 pairs
EN↔ES
EN↔DE
EN↔FR
EN↔IT
EN↔PT
Miami 10h
1
AssemblyAI Universal-3.6 Pro Realtime
7.20%
9.59%
6.13%
7.23%
6.03%
7.00%
24.28%
2
AssemblyAI Universal-3.5 Pro Realtime
8.65%
10.95%
7.84%
8.54%
7.21%
8.71%
27.31%
3
AWS Transcribe Streaming
9.13%
11.32%
8.12%
9.33%
7.05%
9.81%
27.96%
4
Google Gemini 3.5 Transcribe Live
10.16%
14.40%
8.29%
10.05%
8.25%
9.79%
40.69%
5
Reson8 Reson8
10.67%
12.58%
9.73%
12.60%
8.77%
9.67%
24.48%
6
Soniox Realtime
11.26%
12.12%
10.81%
12.15%
10.04%
11.18%
21.93%
7
ElevenLabs Scribe v2 Realtime
13.01%
14.27%
12.29%
13.16%
11.53%
13.79%
27.09%
8
Meta Muse Voice Transcribe
13.34%
13.34%
12.31%
14.64%
11.56%
14.85%
23.10%
9
Smallest Pulse
14.56%
18.52%
13.14%
13.61%
10.79%
16.73%
47.17%
10
Azure Realtime STT
14.60%
17.25%
13.33%
14.86%
12.52%
15.03%
34.17%
11
Gladia Solaria
17.09%
18.39%
18.10%
18.73%
14.28%
15.93%
34.49%
12
Deepgram Nova-3 Multi
17.13%
16.66%
18.01%
18.52%
15.17%
17.29%
29.91%
13
Mistral Voxtral Mini Realtime
18.46%
18.37%
19.11%
20.79%
16.61%
17.43%
34.35%
14
xAI Grok Streaming
20.61%
21.73%
20.06%
21.69%
18.79%
20.79%
34.20%
15
Zoom Scribe (Live)
32.48%
26.66%
35.09%
32.56%
32.09%
36.02%
37.87%
16
OpenAI GPT-4o mini Transcribe
33.27%
35.28%
33.06%
34.22%
31.73%
32.07%
45.01%
17
OpenAI GPT-4o Transcribe
35.46%
39.66%
34.70%
35.78%
33.17%
34.00%
61.27%
18
OpenAI GPT Realtime
37.35%
39.34%
35.46%
37.87%
36.30%
37.77%
53.31%
19
Speechmatics Enhanced Realtime
38.23%
33.75%
33.78%
35.98%
45.39%
42.23%
36.04%
20
Speechmatics Standard Realtime
42.86%
37.60%
44.12%
42.99%
45.11%
44.49%
40.95%
21
Cartesia Ink-Whisper EN
50.52%
46.90%
51.50%
53.10%
53.17%
47.95%
43.72%
22
Deepgram Flux EN
52.86%
52.36%
52.33%
56.46%
53.39%
49.76%
45.55%
23
Deepgram Nova-3 EN
54.71%
54.39%
54.61%
58.15%
56.24%
50.15%
43.85%
Best Others The best value in each column is filled blue.
Figure 7Five-pair mean WER.normalized word error rate · runs 2026-08-18 → 2026-09-23 · one row per model, most recent run
Audio that alternates between English and the second language mid-stream, built from public read-speech corpora in short, long and split-long clips so the switch falls at different points.
◆Miami, 10-hour selection 21 recordings
A ten-hour selection of the Miami Spanish-English conversation, with references produced independently by a third-party transcription service.
○public, open source ◆public, licensed ●private, ours
07Diarization
Speaker attribution on CALLHOME, AMI, DiPCo and NOTSOFAR, for systems that offer streaming diarization. Diarization error rate is the share of audio time given to the wrong speaker, missed, or wrongly marked as speech.
Table 15Diarization error rate and cpWER. Reson8 caps speaker count at 4, its product maximum. That fits CALLHOME, AMI and DiPCo, whose references never exceed four speakers, and costs it on NOTSOFAR, which runs to seven, so its per-corpus figures are not directly comparable with systems that estimate speaker count freely.diarization error rate (DER) and cpWER · runs 2026-08-20 → 2026-09-24 · one row per model, most recent run
Diarization error rate and cpWER. Reson8 caps speaker count at 4, its product maximum. That fits CALLHOME, AMI and DiPCo, whose references never exceed four speakers, and costs it on NOTSOFAR, which runs to seven, so its per-corpus figures are not directly comparable with systems that estimate speaker count freely.
#
System model / snapshot
DER mean 4 sets
CALLHOME
AMI
DiPCo
NOTSOFAR
cpWER mean
CALLHOME
AMI
DiPCo
NOTSOFAR
1
Meta Muse Voice Transcribe
14.51%
16.44%
19.45%
14.81%
7.34%
24.23%
12.54%
22.66%
26.65%
35.07%
2
AssemblyAI Universal-3.5 Pro Realtime
15.96%
13.87%
18.31%
21.29%
10.37%
30.59%
19.76%
28.00%
34.00%
40.61%
3
AssemblyAI Universal-3.6 Pro Realtime
16.68%
13.74%
18.79%
22.10%
12.09%
30.92%
19.21%
28.16%
34.20%
42.11%
4
Reson8 Reson8
16.72%
15.99%
15.93%
15.98%
18.98%
34.32%
22.34%
28.21%
34.61%
52.11%
5
Azure Realtime STT
20.43%
20.21%
18.00%
24.30%
19.19%
40.96%
34.68%
33.63%
41.45%
54.06%
6
Speechmatics Enhanced Realtime
22.16%
14.58%
19.96%
28.37%
25.74%
37.14%
24.69%
23.01%
42.59%
58.25%
7
xAI Grok Streaming
25.68%
15.59%
31.90%
35.28%
19.96%
41.81%
25.77%
37.95%
44.72%
58.82%
8
Speechmatics Standard Realtime
25.71%
25.43%
21.27%
24.68%
31.47%
48.86%
47.84%
33.66%
43.43%
70.49%
9
AWS Transcribe Streaming
27.48%
27.89%
26.16%
32.09%
23.79%
46.07%
41.66%
38.72%
47.91%
55.98%
10
Soniox Realtime
31.77%
19.83%
46.27%
27.23%
33.76%
42.78%
17.20%
55.47%
39.56%
58.88%
11
Deepgram Nova-3
34.01%
30.46%
33.86%
45.35%
26.37%
56.09%
52.17%
43.99%
69.88%
58.30%
12
Smallest Pulse
41.35%
27.16%
59.41%
50.98%
27.87%
53.34%
31.65%
65.89%
62.56%
53.26%
Best Others The best value in each column is filled blue.
Data
◆CALLHOME 40 calls
Public CALLHOME telephone calls between friends and family: unscripted, fast, heavily overlapping conversation over a real phone line.
○AMI 24 meetings
Public AMI Meeting Corpus: small groups of colleagues in real design and project meetings, in instrumented rooms.
○NOTSOFAR 129 meetings
Public NOTSOFAR corpus: real multi-party meetings on a single distant microphone in ordinary meeting rooms.
○DiPCo 10 sessions
Public Dinner Party Corpus: groups talking over dinner on far-field microphones, constant overlap and background clatter.
○public, open source ◆public, licensed ●private, ours
08Independent evaluations
Third-party results on their own corpora and pipelines, reproduced for comparison.
Artificial Analysis — Speech to Text (Streaming)
AA-WER v2.2 (May 2026): 50% AA-AgentTalk, 25% VoxPopuli, 25% Earnings22. End of speech is detected with SileroVAD; latency runs from there to the final transcript. Price is per 1,000 minutes of audio, normalized across billing models. The leaderboard publishes no “last updated” date. The two AssemblyAI configurations shown are Universal-3.5 Pro Realtime. Universal-3.6 Pro has not been scored here yet; this section will be updated once the source publishes it.
Table 1630 of the 31 systems on the Artificial Analysis streaming leaderboard, at published ranks; our earlier-generation model is omitted.source = https://artificialanalysis.ai/speech-to-text/streaming · read 2026-09-09 · price shown for every system the source publishes one for
30 of the 31 systems on the Artificial Analysis streaming leaderboard, at published ranks; our earlier-generation model is omitted.
#
System
AA-WER index %
t final s
Price $/1k min
1
Muse Voice Transcribe (Meta)
3.06%
0.163 s
3.0
2
Cartesia Ink-2 (semantic)
3.36%
0.431 s
4.0
3
ElevenLabs Scribe v2 Realtime
3.59%
0.141 s
6.5
4
Qwen3 ASR Flash Realtime
3.73%
0.476 s
5.4
5
GPT Live Transcribe
3.92%
0.812 s
17.0
6
Grok STT Streaming
3.93%
0.373 s
3.3
7
Gemini 3.5 Transcribe Live
4.00%
0.395 s
9.0
8
AssemblyAI U3.5 Realtime Pro — min latency
4.02%
0.191 s
7.5
9
Cartesia Ink-2 (external endpoint)
4.02%
0.067 s
4.0
10
AssemblyAI U3.5 Realtime Pro — max accuracy
4.03%
0.202 s
7.5
12
Inworld STT 1
4.18%
0.071 s
1.4
13
Soniox v4
4.49%
0.062 s
2.0
14
Soniox v5
4.50%
0.054 s
2.0
15
Google Chirp 3 Streaming
4.80%
1.276 s
16.0
16
OpenAI GPT Realtime (Whisper)
4.89%
0.688 s
17.0
17
Smallest Pulse
4.98%
0.132 s
4.0
18
Voxtral Mini Realtime
5.24%
0.682 s
6.0
19
Azure Speech
5.25%
0.625 s
16.7
20
Nemotron 3 ASR (1120 ms)
5.36%
0.418 s
—
21
Amazon Transcribe
5.79%
0.620 s
24.0
22
Nemotron 3 ASR (560 ms)
5.96%
0.252 s
—
23
Deepgram Nova-3 Realtime
6.59%
0.066 s
4.8
24
Gradium
6.61%
0.435 s
12.0
25
Nemotron 3 ASR (80 ms, Together)
7.03%
0.070 s
—
26
Deepgram Flux
7.39%
0.021 s
6.5
27
Nemotron 3 ASR (160 ms)
7.61%
0.098 s
—
28
Gladia Solaria
7.82%
1.523 s
12.5
29
Speechmatics Realtime Enhanced
8.05%
0.317 s
17.5
30
Nemotron 3 ASR (80 ms)
8.38%
0.073 s
—
31
Rev.ai
10.14%
0.697 s
3.3
Best Others The best value in each column is filled blue.
Table 17The three sets in the Artificial Analysis index. The source ranks our two configurations separately; both are shown beside each set’s leading system.source = https://artificialanalysis.ai/speech-to-text/streaming · read 2026-09-09
The three sets in the Artificial Analysis index. The source ranks our two configurations separately; both are shown beside each set’s leading system.
Set
Our configuration
WER
Rank
Leading system on the set
Its WER
VoxPopuli
U3.5 RT Pro — min latency
2.06%
3 of 31
Gemini 3.5 Transcribe Live
1.76%
VoxPopuli
U3.5 RT Pro — max accuracy
2.09%
4 of 31
Gemini 3.5 Transcribe Live
1.76%
AA-AgentTalk
U3.5 RT Pro — min latency
3.53%
9 of 31
Qwen3 ASR Flash Realtime
2.49%
AA-AgentTalk
U3.5 RT Pro — max accuracy
3.48%
7 of 31
Qwen3 ASR Flash Realtime
2.49%
Earnings22
U3.5 RT Pro — min latency
6.96%
11 of 31
Muse Voice Transcribe
4.31%
Earnings22
U3.5 RT Pro — max accuracy
7.08%
12 of 31
Muse Voice Transcribe
4.31%
1 ElevenLabs Scribe v2 Realtime · 0.141 s · 3.59%
2 Gemini 3.5 Transcribe Live · 0.395 s · 4.00%
3 Google Chirp 3 Streaming · 1.276 s · 4.80%
4 OpenAI GPT Realtime (Whisper) · 0.688 s · 4.89%
5 Nemotron 3 ASR (560 ms) · 0.252 s · 5.96%
Figure 8AA-WER against time to final transcript. The dashed line is the accuracy–latency frontier.source = https://artificialanalysis.ai/speech-to-text/streaming
Coval — Live STT Leaderboard
7-day window, read 2026-09-28 (snapshot generation 1045). The source’s default view is a rolling 24 hours; this is the wider window. WER is pooled over seven datasets — PipeCat production speech (897 clips) and six WildASR conditions (accents 845 clips, clean, clipping, far-field, noise gaps and phone codec at 284 each). Each paired WildASR condition reuses the clean baseline’s utterance, so a difference isolates the condition. Latency is the per-model average. The PipeCat references are model-generated rather than human-verified, which the source notes makes that set better for comparing models than for reading an absolute error rate. Universal-3.6 Pro’s three figures were supplied ahead of the source’s own publication of that model; its placings are computed by inserting them into the 2026-09-28 ranking rather than read from it, and should be re-read once the source publishes. Coval sells voice-agent evaluation tooling.
Table 18Selected rows of 30. Shown as published, except Universal-3.6 Pro, whose figures precede the source’s publication of that model. Three AssemblyAI models are listed.source = https://benchmarks.coval.ai/stt · read 2026-09-28 · 7-day window
Selected rows of 30. Shown as published, except Universal-3.6 Pro, whose figures precede the source’s publication of that model. Three AssemblyAI models are listed.
System model / snapshot
WER pooled %
AssemblyAI Universal 3.6 Pro rank 1 of 30
2.30%
AssemblyAI Universal 3.5 Pro rank 2 of 30
2.48%
Nari Labs Qwen3 ASR Fast rank 3 of 30
2.72%
Reson8 resonant-1 rank 4 of 30
2.76%
Baseten Qwen3 ASR 1.7b rank 5 of 30
2.78%
Gemini 3.5 Transcribe Live rank 6 of 30
3.01%
ElevenLabs Scribe v2 Realtime rank 17 of 30
4.47%
Soniox STT RT v5 rank 19 of 30
4.80%
Deepgram Flux rank 22 of 30
5.77%
AssemblyAI Universal Streaming rank 23 of 30
5.78%
Best Others The best value in each column is filled blue.
Figure 9Universal-3.5 Pro’s WER by recording condition — the source publishes no per-condition breakdown for Universal-3.6 Pro. The dashed line is Universal-3.5 Pro’s own pooled value from Table 18.source = https://benchmarks.coval.ai/stt
Pipecat / Daily — STT benchmark
1,000 samples from the smart-turn-data-v3.1-train dataset; the source does not describe the set as English. Pooled WER is total errors over total reference words, judged semantically; TTFS is the median and P95 from when the speaker stops to the final transcription segment. Reproduced with the source’s own model names; ranked here by Pooled WER, where the source now lists rows alphabetically by vendor. Read from the repository README (last changed 2026-09-18). The source also publishes a per-sample WER mean, a transcript count, a perfect-transcript count and a P99, which are not carried here.
Table 19Shown as published. Three AssemblyAI entries are listed; the source marks universal-3-6-pro as the current one and the other two as superseded.source = https://github.com/pipecat-ai/stt-benchmark · read 2026-09-28
Shown as published. Three AssemblyAI entries are listed; the source marks universal-3-6-pro as the current one and the other two as superseded.
#
System
Pooled WER %
TTFS median ms
TTFS p95 ms
1
Meta muse-voice-transcribe-1.0
0.83%
392
1292
2
AssemblyAI universal-3-6-pro
0.96%
307
401
3
Speechmatics linden-1
1.05%
369
438
4
Speechmatics
1.07%
495
676
5
Azure
1.18%
1016
1345
6
AssemblyAI universal-3-5-pro
1.22%
282
354
7
Cartesia ink-2
1.25%
299
328
8
Soniox stt-rt-v5
1.27%
260
305
9
Soniox stt-rt-v4
1.29%
249
281
10
Deepgram nova-3-general
1.62%
247
298
11
AWS
1.75%
1136
1527
12
NVIDIA Nemotron 3.0 ASR (en)
1.95%
221
238
13
Google gemini-3.5-transcribe-live
2.24%
458
532
14
Smallest AI pulse
2.37%
398
533
15
OpenAI gpt-realtime-whisper
2.73%
740
878
16
Google latest-long
2.85%
878
1155
17
AssemblyAI universal-streaming-english
3.02%
256
362
18
OpenAI gpt-4o-transcribe
3.06%
637
965
19
ElevenLabs scribe_v2_realtime
3.12%
281
348
20
Gradium default
3.71%
570
596
21
Cartesia ink-whisper
4.36%
266
364
22
NVIDIA Nemotron 3.5 ASR (multilingual)
4.58%
236
253
23
Mistral voxtral-mini-transcribe-realtime-2602
4.97%
525
973
Best Others The best value in each column is filled blue.
09How we measure
How the results were produced: the harness, the metrics, why public leaderboards can rank the same systems differently, and the limits of this report.
How we run every system
Protocol. Audio is streamed in real time to each system’s public API, default settings, language fixed. One normalizer is applied before scoring. Each streaming model a vendor ships has its own row.
Complete runs only. An average is shown only when the system was scored on every dataset in it; partial rows keep their cells but are not ranked.
n/d, n/s. A 0.00 error rate means no data. WER or CER of 60% or more means the language is unsupported; the cell is excluded from means.
Excluded. Unreleased models, and results on a superseded version of the voice-agent test set.
Systems
Every system with a public streaming API and a current run; a dot marks each suite it was scored on.
Table 20Systems evaluated and the suites each has a current run on.5 suites · runs 2026-08-27 → 2026-09-22 · added this update: Meta, Zoom
Systems evaluated and the suites each has a current run on.
The audio each benchmark uses and the metric it scores. Each results section also lists its own datasets.
Table 21Test data by benchmark. Files are scored recordings or clips; multilingual, code-switching and judge-scored sets are scored per language or label, so no single count applies. Public corpora are named, licensed sets described.
Test data by benchmark. Files are scored recordings or clips; multilingual, code-switching and judge-scored sets are scored per language or label, so no single count applies. Public corpora are named, licensed sets described.
Suite
Audio
Files
Scored on
Voice agents
Scripted caller scenarios voiced by real speakers, recorded in three noise environments, three utterance-length bands and 12 speaker groups
Short-form English audio, with entities graded per label by an LLM judge
—
Entity error (judge)
Non-speech
Hold music, line noise and silence with no speech
2,950
Response rate
Multilingual
FLEURS and Common Voice test sets in eighteen languages
—
WER (character error rate for Japanese, Chinese)
Code-switching
Five English↔X pairs derived from read speech, plus a bilingual conversational corpus
—
WER
Diarization
CALLHOME calls; AMI and NOTSOFAR meetings; DiPCo dinner parties
203
DER, cpWER
Each test set is marked by how its audio can be obtained: ○ public and open source; ◆ public but licensed or purchasable; ● private, recorded or annotated by us.
How we score
Each metric and how it is computed. The same normalizer and scoring code apply to every system.
WER Word error rate after normalization; lower is better. Character error rate for Japanese, Chinese and Korean.
NEER Normalized entity error rate: the share of entities (names, addresses, numbers, emails, dates, codes, medical and technical terms) transcribed incorrectly.
MER Medical entity error rate over medication and disease terms.
Judge-scored entity error Entity error rate as graded by an LLM judge.
DER / cpWER Diarization error rate, and concatenated minimum-permutation WER.
Non-speech response rate Share of non-speech clips for which the system produces any text.
Streaming harness. Audio is chunked and streamed in real time to each vendor’s endpoint through its documented SDK or WebSocket API, default settings, language fixed. Final transcripts are scored, not partials.
Normalization. One normalizer (case, punctuation, numbers, abbreviations) is applied to references and to every system’s output. Entity metrics are computed on normalized text and, where reported, on formatted text.
Entity error. Entity spans are located in the reference; an entity is an error if any token of its span is wrong in the aligned transcript. The judge-scored variant asks an LLM whether the transcript names the same real-world entity.
Selection rule. Every row is the model’s most recent production run. The overview grid keeps each vendor’s best model per column and names it on hover.
Provenance. Each figure carries its metric and run-date span. Public test sets are named; licensed sets are described.
Why public leaderboards and ours differ
We build for voice agents, clinical documentation, meetings, telephony, and multilingual and code-switched speech, so the test sets cover that range of speakers, accents, languages and acoustic conditions and score what matters in each. Public leaderboards usually ask a narrower question, typically one word-error-rate figure on a few corpora, so a model can rank differently on the two.
Table 22What each benchmark measures: this report and the public leaderboards reproduced in §9.
What each benchmark measures: this report and the public leaderboards reproduced in §9.
Benchmark
Audio
Scores
Run by
References
This report
Voice-agent calls in three noise environments, clinical consultations and dictation, accented and long-form English, eighteen languages, code-switched speech, meetings and telephone calls, non-speech audio
One WER index over the three sets; time to final transcript
Artificial Analysis
Published with the sets
Coval
Short clips under seven recording conditions (clean, accent, far-field, phone codec, reverb…), 7-day window
Pooled WER; time to first token; time to final segment
Coval
Model-generated
Pipecat / Daily
1,000 English samples
Semantic WER; median and P95 time to final
Daily
Published with the set
The Artificial Analysis index is one WER over three sets, half of it AA-AgentTalk, with no entity, noise, medical, multilingual, code-switching or diarization component; its latency runs from a voice-activity detector’s end of speech to the final transcript, which this report does not measure. Coval scores against model-generated references. Neither substitutes for the other or for this report.
Caveats
We ran it. AssemblyAI produced the results in §1–§8. Competitors ran on default settings and may do better tuned. §9 has independent results.
Snapshots differ. Runs span 2026-08-18 to 2026-09-24 (Table 20). A vendor that updated its model within that window may appear on the older version.
Coverage is uneven. Not every system ran every set; means are withheld rather than computed on partial data.
Reference quality. TED-LIUM references omit sponsor read-outs. Some independent leaderboards (Coval) use model-generated references.
AssemblyAI Research · Realtime benchmarks technical report · data as of 2026-09-16
Benchmarks
Run your own benchmark
Testing on your own audio? Want to do it right? Read our benchmarking guide in the docs.
Benchmark audio was evaluated with production model settings. No custom model tuning or prompt engineering was applied.
Transcription outputs were normalized using the Whisper text normalizer before WER computation. This removes formatting differences (casing, punctuation, number formats) so that scores reflect transcription quality, not output styling.
We evaluate across benchmark datasets spanning general speech, entity recognition, diarization, multilingual audio, and code switching.
Word error rate
Selected speech sets covering synthetic medical, accented English, general speech, and webinar audio.