The first API built for dictation. Your users speak, and it returns text that is ready to send: filler gone, self-corrections resolved, names spelled right.
The first API built for dictation. Your users speak, and it returns text that is ready to send: filler gone, self-corrections resolved, names spelled right.
Benchmarks for Universal-3.5 Pro across pre-recorded audio and Universal-3.5 Pro Realtime across realtime audio.
Word error rate
WER is calculated as (substitutions + insertions + deletions) / total words in the reference transcript. It is the standard metric for evaluating speech-to-text accuracy.
Average normalized WER across selected datasets
*lower is better*
4.35%
5.24%
5.34%
5.50%
5.87%
5.91%
6.47%
6.66%
7.02%
7.16%
17.39%
AssemblyAI Universal-3.5 Pro
Mistral Voxtral Mini
OpenAI GPT-4o Transcribe
Cohere Transcribe
ElevenLabs Scribe V2
Qwen3 ASR
Gladia
Deepgram Nova-3
Azure Batch
Grok
Soniox
WER by dataset expand_more
Dataset
AssemblyAI Universal-3.5 Pro
ElevenLabs Scribe V2
Mistral Voxtral Mini
OpenAI GPT-4o Transcribe
Cohere Transcribe
Qwen3 ASR
Gladia
Deepgram Nova-3
Azure Batch
Grok
Soniox
Synthetic medical
0.33%
0.41%
1.25%
0.55%
1.33%
1.15%
1.23%
0.51%
1.55%
1.38%
0.72%
Accented English (India)
5.19%
5.92%
6.44%
6.49%
6.39%
6.61%
6.78%
7.77%
8.04%
6.60%
53.48%
General speech
6.24%
7.19%
6.64%
7.45%
8.35%
7.01%
11.08%
8.81%
8.42%
9.76%
7.55%
Webinar speech
5.63%
9.96%
6.65%
6.87%
5.91%
8.85%
6.79%
9.55%
10.07%
10.90%
7.79%
Average
4.35%
5.87%
5.24%
5.34%
5.50%
5.91%
6.47%
6.66%
7.02%
7.16%
17.39%
Missed entity rate
Missed entity rate measures errors on named entity categories including names, organizations, locations, medical terms, money, occupations, temporal expressions, and URLs. Unlike aggregate WER, it isolates the tokens that carry the most semantic weight in downstream applications.
Missed entity rate by provider
*lower is better*
Name
Organization
Location
Medical
Money
Occupation
Temporal
Url
24.1
19.0
8.7
13.9
43.7
9.0
16.3
54.4
25.9
20.8
10.0
19.8
39.2
9.4
6.2
65.0
29.2
23.4
12.2
16.1
37.3
13.0
13.8
65.5
20.9
16.6
11.5
10.8
76.9
9.1
20.4
47.1
25.6
19.3
13.6
15.7
76.8
9.7
19.4
54.8
36.1
31.6
24.3
31.1
80.4
10.6
29.7
98.1↑
AssemblyAI Universal-3.5 Pro
Mistral Voxtral Mini
OpenAI GPT-4o Transcribe
ElevenLabs Scribe V2
Deepgram Nova-3
NVIDIA Canary 1B
MER by entity type expand_more
Entity type
AssemblyAI Universal-3.5 Pro
Mistral Voxtral Mini
OpenAI GPT-4o Transcribe
ElevenLabs Scribe V2
Deepgram Nova-3
NVIDIA Canary 1B
Name
24.08%
25.89%
29.23%
20.87%
25.58%
36.08%
Organization
18.99%
20.80%
23.40%
16.57%
19.25%
31.60%
Location
8.72%
9.98%
12.21%
11.54%
13.57%
24.27%
Medical
13.87%
19.78%
16.07%
10.80%
15.69%
31.07%
Money
43.70%
39.22%
37.30%
76.88%
76.82%
80.40%
Occupation
9.01%
9.43%
13.01%
9.12%
9.67%
10.59%
Temporal
16.33%
6.21%
13.78%
20.42%
19.42%
29.70%
Url
54.37%
65.05%
65.53%
47.09%
54.79%
98.06%
Diarization
Diarization segments audio by speaker. We report cpWER (concatenated minimum-permutation word error rate), which jointly evaluates transcription and speaker assignment.
Average cpWER across DiPCo, CALLHOME, NOTSOFAR, and AMI
*lower is better*
30.17%
30.35%
35.26%
36.60%
36.88%
37.52%
37.93%
42.58%
50.64%
114.81%
AssemblyAI Universal-3.5 Pro
Azure
ElevenLabs Scribe V2
Speechmatics
Gladia
Mistral Voxtral
Deepgram
Grok
Google
Soniox
cpWER by dataset expand_more
Dataset
AssemblyAI Universal-3.5 Pro
Azure
ElevenLabs Scribe V2
Speechmatics
Gladia
Mistral Voxtral
Deepgram
Grok
Google
Soniox
DiPCo
33.48%
33.23%
48.78%
36.88%
41.90%
38.04%
38.31%
41.60%
56.40%
126.82%
CALLHOME dev/eval
17.18%
20.28%
14.82%
21.41%
20.06%
19.50%
23.70%
36.85%
29.89%
103.85%
CALLHOME train
17.78%
20.14%
16.18%
20.61%
21.02%
19.36%
25.47%
33.68%
31.12%
102.55%
NOTSOFAR test
37.02%
35.68%
44.37%
56.53%
47.81%
51.91%
48.24%
52.35%
63.83%
111.86%
NOTSOFAR dev
48.22%
45.38%
51.29%
63.38%
56.12%
61.42%
62.55%
58.92%
75.63%
116.41%
AMI
27.36%
27.39%
36.14%
20.82%
34.34%
34.86%
29.28%
32.10%
46.96%
127.36%
Average
30.17%
30.35%
35.26%
36.60%
36.88%
37.52%
37.93%
42.58%
50.64%
114.81%
Multilingual
Global WER across the evaluated multilingual benchmark suite, with language-level breakdowns where available.
Global multilingual WER
*lower is better*
7.84%
8.22%
9.52%
10.50%
11.13%
13.36%
14.39%
15.71%
AssemblyAI Universal-3.5 Pro
Speechmatics Enhanced
OpenAI GPT-4o Transcribe
Mistral Voxtral Mini
ElevenLabs Scribe V2
Cohere Transcribe
OpenAI Whisper-1
Deepgram Nova-3
Global WER and language breakdown expand_more
Model
Global WER
Evaluated languages
German
Spanish
French
Italian
Portuguese
AssemblyAI Universal-3.5 Pro
7.84%
5/20
9.07%
6.57%
10.17%
6.74%
6.65%
Speechmatics Enhanced
8.22%
5/20
10.30%
5.84%
10.09%
7.88%
7.01%
OpenAI GPT-4o Transcribe
9.52%
15/20
—
—
—
—
—
Mistral Voxtral Mini
10.50%
10/20
—
—
—
—
—
ElevenLabs Scribe V2
11.13%
19/20
—
—
—
—
—
Cohere Transcribe
13.36%
—
—
—
—
—
—
OpenAI Whisper-1
14.39%
—
—
—
—
—
—
Deepgram Nova-3
15.71%
—
—
—
—
—
—
Code switching
Code-switching benchmarks test Common Voice-derived English paired with German, Spanish, French, Italian, and Portuguese.
Average code-switching WER
*lower is better*
7.60%
10.07%
12.42%
12.83%
21.88%
22.72%
22.72%
28.13%
36.05%
39.38%
42.71%
46.19%
46.39%
51.24%
52.40%
AssemblyAI Universal-3.5 Pro
ElevenLabs Scribe V2
Soniox
Deepgram Nova-3
Qwen3 ASR
Mistral Voxtral Mini
VibeVoice
Cohere Transcribe
Speechmatics
Grok
OpenAI GPT-4o Transcribe
Gladia
AWS Transcribe
Azure Batch
Gemini
WER by language pair and condition
*lower is better*
en-es long
en-de short
en-fr short
en-it short
en-pt short
5.9
5.9
7.7
6.9
11.7
5.5
10.6
11.4
11.6
11.3
8.5
11.8
14.8
11.6
15.5
10.1
12.7
14.2
12.5
14.6
9.1
24.9
27.0
22.0
26.4
7.8
27.1
29.6
27.1
21.9
11.4
24.4
27.4
22.8
27.6
28.8
27.4
32.5
27.3
24.6
27.3
32.8
35.8
43.0
41.5
12.5
46.7
47.3
42.5
47.9
34.9
42.9
45.5
46.6
43.7
49.7
47.1
42.4
49.4
42.3
43.0
46.0
51.2
47.7
44.1
48.5
55.6↑
54.7
50.8
46.5
47.5
53.2
56.9↑
55.4↑
49.0
AssemblyAI Universal-3.5 Pro
ElevenLabs Scribe V2
Soniox
Deepgram Nova-3
Qwen3 ASR
Mistral Voxtral Mini
VibeVoice
Cohere Transcribe
Speechmatics
Grok
OpenAI
Gladia
AWS Transcribe
Azure Batch
Gemini
Price-performance
Choose a metric to see how each provider's accuracy compares at their price point.
Price per hour against the selected metric
*bottom-left is best*
Pricing from Artificial Analysis.
Demo
Listen to the difference
Play the preloaded sample and compare the transcripts side by side.
Medication name changes
Ramipril becomes a different medication, which can change the patient record.
Just use my personal email. It's jonjay@freemail.com. That's J-O-N-J-A-Y@freemail.com. So jonjay@freemail.com.
AssemblyAI
Just use my personal email. It's jonjay@freemail.com. That's J-O-N-J-A-Y@freemail.com. So jonjay@freemail.com.
Competitor
Just use my personal email. It's johnjay@female.com. That's jonjay@freemail.com. So johnj@freemail.com.
Multiple speakers collapse into one label
Distinct turns from Chris, Sally, Donald, and Jenny are merged under speaker 0.
0:00 / 0:00
import assemblyai as aai
aai.settings.api_key = "YOUR_API_KEY"
config = aai.TranscriptionConfig(
speech_model="universal-3.5-pro",
speaker_labels=True,
)
transcript = aai.Transcriber().transcribe(
"meeting.mp3", config=config
)
for u in transcript.utterances:
print(f"Speaker {u.speaker}: {u.text}")
Truth
Chris:yeah
Sally:mmm
Donald:i was thinking in about uh four months they're having the county fair perhaps you can set a booth up at the county fair
Chris:that's also a great idea um we would have to coordinate that with the city would anyone um be willing to take care of that
Jenny:yes i can take care of that and talk to the city coordinate it with them i think it's a great idea that we have it at the fair like uh donald mentioned
Chris:um ok so uh what else uh does anybody have any ideas of what
AssemblyAI
A:Yeah, I was thinking about, uh, 4 months they're having the county fair. Perhaps we could set a booth up at the county fair.
B:That's also a great idea. Um, we would have to coordinate that with the city. Would anyone be willing to take care of that?
C:Yes, I can take care of that and talk to the city, coordinate it with them. I think it's a great idea that we have it at the fair like Donald mentioned.
B:Um, okay, so, uh, what else? Uh, does anybody have any ideas of what
Competitor
Deepgram diarization
0:Yeah. Mhmm. I was thinking that in about four months, they're having the county fair. Perhaps you could set a booth up at the county fair. That's also a great idea. We would have to coordinate that with the city. Would anyone be willing to take care of that? Yes. I can take care of that and talk to the city, coordinate it with them. I think it's a great idea that we have it at the fair, like Donald mentioned.
0:Okay. So what else? Does anybody have any ideas of what
A band name changes meaning
NOFX becomes No Effect, turning a name into a phrase.
Technical reportRealtime speech-to-textdata as of 2026-09-09v2026.09
Realtime speech recognition, measured against the field
Universal-3.5 Pro Realtime and 24 competitor models on 6 evaluation suites: same audio, same scoring pipeline.
Model under test Universal-3.5 Pro Realtime (production endpoint)Evaluation AssemblyAI ResearchRivals run via each vendor’s public streaming API, default settingsSnapshot 2026-09-09
Universal-3.5 Pro Realtime has the lowest entity error rate on the voice-agent suite: 14.78% against 15.95% for the next system, a 7% lead. It also leads on judge-scored entity accuracy in general English, speaker diarization, code-switched speech and end-of-turn recall.
Every score is in the tables. §9 reproduces independent evaluations; §10 describes the method.
At a glance
Voice-agent entity error
14.78%
1st of 17
next best ElevenLabs Scribe v2 15.95%
Judge-scored entity error
19.17%
1st of 17
next best ElevenLabs Scribe v2 21.19%
Diarization DER
17.62%
1st of 9
next best Azure 20.43%
Code-switching (5 pairs)
7.82%
1st of 16
next best AWS Transcribe 8.30%
Multilingual (FLEURS, 9 lang.)
7.20%
1st of 13
next best AWS Transcribe 7.37%
01Overview
Each cell is the vendor’s best model on that metric, coloured by rank within the column (solid blue is first). Values are in the tables of §2–§8.
Rank across 11 headline metrics, vendors ordered by mean rank.
Vendor (best model per cell)
VA entity err. NEER %
Entity (judge) %
Medical MER %
End-of-turn F1 %
Endpoint P90 ms
Diarization DER %
Non-speech resp. %
Code-switching WER %
Multi-lingual FLEURS %
EN short-form WER %
EN long-form WER %
meanrank
AssemblyAI
14.8
19.2
5.1
81.9
1622
17.6
0.3
7.8
7.2
3.9
9.5
3.1
Speechmatics
27.8
24.8
4.7
28.1
4618
22.2
0.2
38.1
12.4
4.7
7.7
5.9
ElevenLabs
15.9
21.2
—
85.5
1955
—
0.2
12.4
7.7
—
19.9
6.0
AWS
22.8
30.7
—
77.8
965
27.5
1.3
8.3
7.4
—
8.0
6.0
Deepgram
33.8
30.7
13.2
76.6
748
34.0
0.2
16.5
—
6.5
9.1
6.6
Mistral
23.1
23.4
12.3
26.7
3991
—
0.0
17.7
9.7
6.6
22.7
7.2
Azure
—
50.2
14.4
—
—
20.4
0.2
13.8
9.8
5.5
8.7
7.3
Google
21.6
—
8.6
84.1
1482
—
—
—
—
—
—
7.5
Smallest
46.8
40.7
21.4
74.0
1247
41.4
1.4
13.1
12.0
5.4
9.0
7.6
xAI
26.9
36.1
15.7
70.7
952
25.7
1.2
—
—
5.8
9.0
7.7
Reson8
—
25.3
15.0
—
—
—
0.7
—
—
—
6.9
8.8
OpenAI
36.7
25.9
—
56.8
1083
—
0.1
32.8
10.4
—
—
9.0
Soniox
20.4
—
—
24.0
11625
31.5
—
10.8
—
—
—
10.2
Cartesia
35.9
29.8
—
60.9
525
—
3.1
50.7
16.8
—
11.1
10.3
Gladia
—
28.0
—
—
—
—
4.9
16.3
14.1
—
11.7
12.2
BestWorstCells are coloured by rank within each column (solid blue is best, white is worst).
Figure 1Rank across 11 headline metrics, vendors ordered by mean rank.source = the tables below; each column’s provenance is given with its table
Position in the field
Voice-agent entity error14.78% · 1st of 17 systems · runner-up ElevenLabs Scribe v2 15.95% −7.3%
Judge-scored entity error19.17% · 1st of 17 systems · runner-up ElevenLabs Scribe v2 21.19% −9.5%
Medical entity error5.10% · 2nd of 11 systems · leader Speechmatics Enhanced 4.66% +9.4%
End-of-turn F181.94% · 3rd of 17 systems · leader ElevenLabs Scribe v2 85.55% −3.6 pt
Endpointing P90 latency1622 ms · 11th of 17 systems · leader Cartesia Ink-Whisper 525 ms +1097 ms
Diarization DER17.62% · 1st of 9 systems · runner-up Azure 20.43% −13.8%
Non-speech response rate0.27% · 8th of 17 systems · leader Mistral Voxtral Mini 0.03% +0.2 pt
Code-switching (5 pairs)7.82% · 1st of 16 systems · runner-up AWS Transcribe 8.30% −5.8%
Multilingual (FLEURS, 9 lang.)7.20% · 1st of 13 systems · runner-up AWS Transcribe 7.37% −2.3%
English short-form WER3.87% · 1st of 9 systems · runner-up Speechmatics Enhanced 4.69% −17.5%
Long-form English WER9.53% · 10th of 14 systems · leader Reson8 6.86% +38.9%
We lead We trail The margin is blue where we lead, grey where we trail.
One row per headline metric from the grid above: our value, our rank among systems with a complete run, and the margin to the runner-up where we lead or to the leader where we do not. Margins are relative for error rates and absolute for F1, latency, and rates below one percent.
02Voice agents
Conversational audio as a voice agent hears it, across 17 systems, three noise environments, 12 speaker groups and three utterance lengths. The headline metric, entity error (NEER), is how often the names, addresses and numbers an agent acts on come out wrong.
Table 1Entity error rate (%) by entity class. Colour is rank within the column. We lead in 7 of 11 columns.normalized entity error rate (NEER) · runs 2026-08-27 → 2026-09-08 · one row per model, most recent run · corpus value = mean of the three natural-noise environments
Entity error rate (%) by entity class. Colour is rank within the column. We lead in 7 of 11 columns.
System
Overall
Names
Places
Addresses
Phone
Email
Dates
Numbers
Codes/IDs
Medical
Technical
AssemblyAI Universal-3.5 Pro
14.8
11.1
6.1
27.9
2.8
15.7
21.1
15.4
10.4
19.1
26.6
ElevenLabs Scribe v2
15.9
7.3
8.8
29.6
3.3
17.6
27.9
33.6
12.8
12.3
19.9
Soniox Realtime
20.4
16.4
8.1
29.3
3.0
20.8
32.1
26.8
8.6
20.8
32.4
Google Gemini 3.5 Live
21.6
16.0
9.8
32.8
7.8
23.3
33.4
22.3
17.6
22.3
33.0
AWS Transcribe
22.8
21.8
7.4
30.2
3.7
21.0
30.6
24.5
12.8
21.5
40.0
Mistral Voxtral Mini
23.1
18.2
10.1
40.3
8.3
25.9
31.4
21.4
16.1
32.6
37.6
xAI Grok Streaming
26.9
19.0
12.8
85.0
13.7
29.8
24.6
38.6
26.4
37.0
42.5
Speechmatics Enhanced
27.8
24.4
7.0
37.5
85.6
22.7
33.9
21.1
22.9
18.5
38.1
Deepgram Nova-3
33.8
25.5
11.3
31.0
91.7
33.4
47.8
23.1
32.6
30.0
51.1
Cartesia Ink-Whisper
35.9
25.1
16.3
61.6
59.2
47.3
36.0
28.4
55.5
45.1
45.5
Speechmatics Standard
36.3
31.1
9.5
49.4
87.8
56.1
35.4
23.6
46.7
29.8
42.2
OpenAI GPT-4o Transcribe
36.7
28.2
18.3
56.2
44.1
49.6
36.1
34.6
56.3
38.0
45.7
OpenAI GPT-4o mini Transcribe
37.3
29.6
19.9
58.2
41.5
51.7
38.3
31.6
51.2
43.1
46.0
OpenAI GPT Realtime
39.4
32.6
23.3
53.3
39.1
45.6
44.0
36.4
52.8
37.6
45.6
Deepgram Flux EN
46.3
30.6
14.0
94.5
11.6
39.7
44.5
63.0
99.9
33.6
60.5
Smallest Pulse
46.8
30.4
17.0
94.6
13.2
38.7
48.3
55.4
96.5
44.6
53.5
Deepgram Flux Multi
47.6
33.7
14.2
94.6
10.1
36.9
44.4
63.3
99.9
38.8
59.5
BestWorstCells are coloured by rank within each column (solid blue is best, white is worst).
Background noise and utterance length
No background noise Natural low noise Natural heavy noise
Figure 2Entity error rate (%) by background-noise condition, sorted by the no-noise value. We lead in all 3 conditions.normalized entity error rate (NEER) · runs 2026-08-27 → 2026-09-08 · one row per model, most recent run
Listen to the test data
One caller turn per noise environment, quietest first, each with an entity an agent acts on: a name, an account number, an e-mail login, a medication. Words that differ from the human reference are marked.
1 of 8
Voice agentsno noise4 s · AssemblyAI voice-agent corpus — scripted caller scenarios voiced by real speakers
0:00 / 0:00
Four seconds, no noise. A short request where the other systems substitute unrelated words.
Reference
My cable modem keeps rebooting every 30-40 minutes, can you help?
AssemblyAI
My cable modem keeps rebooting every 30 to 40 minutes. Can you help?
ElevenLabs Scribe v2 Realtime
My key won't let them keep changing every 30 to 40 minutes. Can you help?
Soniox Realtime
Mark Kilmer and keeps removing every 30 to 40 minutes. Can you help?
Short utterances Medium Long
Figure 3Entity error rate (%) by utterance length. We lead in all 3 length bands.normalized entity error rate (NEER) · runs 2026-08-27 → 2026-09-08 · one row per model, most recent run
Speaker groups
Table 2Entity error rate (%) by speaker group. Colour is rank within the column. We lead in 10 of 12 groups.normalized entity error rate (NEER) · runs 2026-08-27 → 2026-09-08 · one row per model, most recent run
Entity error rate (%) by speaker group. Colour is rank within the column. We lead in 10 of 12 groups.
System
US Northeast
US South
US Gen. Am.
US AAVE
US Asian-Am.
US other
UK
Canada
Native other
Spanish-acc.
Indian-acc.
Non-nat. other
meanrank
AssemblyAI Universal-3.5 Pro
10.9
12.7
12.4
14.0
13.8
15.8
15.1
11.7
16.5
17.3
21.7
20.2
1.2
ElevenLabs Scribe v2
13.4
15.4
14.2
15.4
14.5
15.6
17.0
13.6
17.0
18.8
20.7
20.9
1.8
Soniox Realtime
17.6
18.4
18.2
19.6
20.0
19.8
21.8
17.5
21.9
24.9
25.7
26.3
3.2
Google Gemini 3.5 Live
18.9
21.4
19.5
21.7
18.5
22.6
23.9
18.4
22.8
26.0
23.3
27.5
4.2
Mistral Voxtral Mini
18.9
21.9
19.8
21.8
21.2
24.1
22.8
17.5
24.9
30.3
30.0
32.4
5.3
AWS Transcribe
19.6
21.7
21.0
22.6
21.5
22.7
22.9
19.8
24.3
26.3
26.9
28.8
5.4
xAI Grok Streaming
21.8
26.9
24.6
26.6
23.9
28.2
28.3
21.9
30.5
29.8
33.5
33.5
7.3
Speechmatics Enhanced
25.7
26.9
25.5
26.3
26.7
28.8
27.9
25.4
29.1
31.7
31.4
33.9
7.7
Deepgram Nova-3
30.7
33.5
31.9
32.9
32.2
33.1
34.5
30.5
35.4
37.9
38.6
39.9
9.3
Cartesia Ink-Whisper
30.3
35.8
32.3
34.5
34.0
38.0
37.3
27.6
38.0
44.5
42.3
46.1
10.5
Speechmatics Standard
32.3
34.9
33.1
36.1
36.2
37.9
36.8
33.6
38.2
42.0
40.9
43.8
11.6
OpenAI GPT-4o Transcribe
31.4
37.3
32.9
35.4
33.2
38.1
38.7
28.4
38.5
46.1
42.4
47.4
11.8
OpenAI GPT-4o mini Transcribe
31.5
37.4
33.0
35.1
34.1
39.1
39.2
29.4
39.4
46.5
44.6
48.9
12.7
OpenAI GPT Realtime
34.9
36.6
34.9
34.5
35.5
37.8
40.9
34.7
43.0
48.8
48.1
51.6
13.2
Deepgram Flux EN
43.9
45.7
43.6
46.5
43.5
44.5
48.6
44.1
48.7
50.5
50.4
52.1
15.3
Smallest Pulse
44.6
44.5
43.4
48.1
44.5
49.7
47.7
40.4
48.7
53.0
53.5
54.7
16.1
Deepgram Flux Multi
45.2
46.7
45.1
48.2
44.4
45.8
49.7
45.1
50.2
52.9
51.9
54.1
16.6
BestWorstCells are coloured by rank within each column (solid blue is best, white is worst).
Turn-taking
Table 3End-of-turn detection and endpointing latency. Higher is better for F1, precision and recall; lower for latency. We lead on recall.end-of-turn detection (F1/precision/recall, %) and endpointing latency (ms) · runs 2026-08-27 → 2026-09-08 · one row per model, most recent run
End-of-turn detection and endpointing latency. Higher is better for F1, precision and recall; lower for latency. We lead on recall.
#
System model / snapshot
F1 %
Precision %
Recall %
P50 ms
P90 ms
P99 ms
Mean ms
1
ElevenLabs Scribe v2 Realtime
85.5
83.3
87.9
1775
1955
2391
1640
2
Google Gemini 3.5 Transcribe Live
84.1
81.5
87.0
1306
1482
2143
1254
3
AssemblyAI Universal-3.5 Pro Realtime
81.9
73.8
92.1
546
1622
2909
742
4
AWS Transcribe Streaming
77.8
69.4
88.4
765
965
1653
728
5
Deepgram Flux Multi
76.6
68.6
86.8
328
748
1479
391
6
Deepgram Flux EN
75.4
68.1
84.5
348
827
1672
402
7
Smallest Pulse
74.0
63.5
88.9
925
1247
2351
954
8
Deepgram Nova-3
71.5
59.3
90.3
397
849
5186
572
9
xAI Grok Streaming
70.7
60.6
85.1
707
952
1554
688
10
Cartesia Ink-Whisper
60.9
48.4
82.5
330
525
747
291
11
OpenAI GPT-4o Transcribe
56.8
44.4
79.0
877
1087
1644
847
12
OpenAI GPT-4o mini Transcribe
55.9
43.9
77.2
880
1083
1645
841
13
Speechmatics Enhanced Realtime
28.1
16.6
91.9
4410
4618
4804
3885
14
Speechmatics Standard Realtime
26.8
15.9
88.5
4543
4752
4906
3947
15
Mistral Voxtral Mini Realtime
26.7
91.6
16.2
816
3991
13085
2061
16
OpenAI GPT Realtime
24.6
84.0
15.0
1740
31091
89510
10124
17
Soniox Realtime
24.0
17.2
39.9
8016
11625
14209
6501
Best Others The best value in each column is filled blue.
1 Deepgram Flux Multi · 748 ms · 76.60%
2 Google Gemini 3.5 Live · 1,482 ms · 84.14%
Figure 4End-of-turn F1 against P90 endpointing latency (log scale). Upper left is better.end-of-turn detection (F1/precision/recall, %) and endpointing latency (ms) · runs 2026-08-27 → 2026-09-08 · one row per model, most recent run
Ablations on our own model
wordboost OFF wordboost ON
0%10%20%30%0%
11.12%
Overall
4.60%
Names
6.04%
Places
28.23%
Addresses
2.62%
Phone
14.53%
Email
21.92%
Dates
15.20%
Numbers
10.01%
Codes/IDs
8.42%
Medical
8.42%
Technical
5.30%
WER
Figure 5Word boost (custom vocabulary) on versus off, same model and corpus. Change in entity error with word boost on: names −58.5%, technical terms −68.4%.normalized entity error rate (NEER) + WER, wordboost OFF vs ON · runs 2026-08-27 → 2026-09-03 · one row per model, most recent run
agent context OFF agent context ON
0%10%20%30%0%
14.13%
Overall
10.16%
Names
5.31%
Places
27.60%
Addresses
2.45%
Phone
15.50%
Email
20.74%
Dates
15.28%
Numbers
10.13%
Codes/IDs
17.46%
Medical
24.56%
Technical
5.61%
WER
Figure 6Agent context on versus off: the agent’s previous turn is passed to the recognizer.normalized entity error rate (NEER) + WER, same model with and without agent context · runs 2026-08-27 → 2026-08-28 · one row per model, most recent run
Word error rate
Table 4WER by noise condition.normalized word error rate · runs 2026-08-27 → 2026-09-08 · one row per model, most recent run
WER by noise condition.
#
System model / snapshot
Mean 3 conditions
No noise
Low noise
Heavy noise
1
AssemblyAI Universal-3.5 Pro Realtime
5.81%
4.63%
5.60%
7.19%
2
Soniox Realtime
6.84%
5.96%
6.86%
7.70%
3
ElevenLabs Scribe v2 Realtime
6.97%
5.59%
6.79%
8.52%
4
Speechmatics Enhanced Realtime
8.14%
6.96%
7.96%
9.49%
5
Google Gemini 3.5 Transcribe Live
8.48%
7.44%
8.49%
9.52%
6
AWS Transcribe Streaming
8.58%
7.27%
8.27%
10.19%
7
Deepgram Nova-3
8.64%
7.53%
8.45%
9.93%
8
Mistral Voxtral Mini Realtime
10.43%
9.25%
10.45%
11.59%
9
Smallest Pulse
10.74%
9.85%
10.94%
11.43%
10
Speechmatics Standard Realtime
12.05%
10.47%
11.77%
13.91%
11
Deepgram Flux EN
13.50%
12.29%
13.36%
14.85%
12
Deepgram Flux Multi
14.26%
12.85%
14.06%
15.87%
13
Cartesia Ink-Whisper
15.17%
13.71%
15.00%
16.81%
14
OpenAI GPT-4o Transcribe
18.37%
17.05%
18.44%
19.63%
15
OpenAI GPT-4o mini Transcribe
19.55%
18.12%
19.77%
20.76%
16
xAI Grok Streaming
20.28%
18.45%
19.89%
22.50%
17
OpenAI GPT Realtime
28.37%
25.45%
28.14%
31.53%
Best Others The best value in each column is filled blue.
Data
●Caller corpus 12,460 clips
All entity classes pooled into one number: how often an agent-critical value (a name, address, number or identifier) comes back wrong.
●Noise environments
Three levels of real noise, recorded rather than mixed in. None: room and handset noise only. Low: incidental noise from a home, office or parked car. Heavy: noise that competes with the voice.
●Utterance length
Short turns (confirmations, single values, one-word answers), medium and long. Short turns give the model almost no context.
●Speaker groups
Fourteen groups: US Northeast, South, General American, AAVE, Asian-American and other; UK; Canada; native speakers overall; Spanish- and Indian-accented and other non-native speakers.
●Entity classes
Names (including letter-by-letter spellings), places, street addresses, phone numbers, email addresses, dates and times, numbers and currency, codes and identifiers, medical and technical terms.
○public, open source ◆public, licensed ●private, ours
03Healthcare
Medical entity error is measured on primary-care consultations (PriMock57), licensed clinical dictation and respiratory-clinic recordings.
Overall
Table 5Medical entity error rate. We lead on primary care.normalized medical entity error rate (MER) · runs 2026-08-18 → 2026-09-02 · one row per model, most recent run
Medical entity error rate. We lead on primary care.
#
System model / snapshot
Mean 3 sets
Primary care PriMock57
Dictation
Respiratory
1
Speechmatics Enhanced Realtime
4.66%
8.13%
3.65%
2.21%
2
AssemblyAI Universal-3.5 Pro Realtime
5.10%
7.18%
5.23%
2.90%
3
Google Gemini 3.5 Transcribe Live
8.57%
15.45%
7.69%
2.56%
4
Mistral Voxtral Mini Realtime
12.26%
12.67%
16.02%
8.09%
5
Deepgram Nova-3
13.22%
19.01%
17.37%
3.29%
6
Speechmatics Standard Realtime
14.24%
18.68%
18.72%
5.32%
7
Deepgram Flux EN
14.28%
20.49%
18.32%
4.02%
8
Azure Realtime STT
14.41%
17.21%
19.11%
6.92%
9
Reson8 Reson8
15.02%
18.49%
21.81%
4.76%
10
xAI Grok Streaming
15.69%
17.95%
19.35%
9.78%
11
Smallest Pulse
21.40%
25.42%
30.77%
8.01%
AWS Transcribe Streaming
incomplete
11.09%
16.49%
—
Cartesia Ink-Whisper
incomplete
24.39%
31.09%
—
ElevenLabs Scribe v2 Realtime
incomplete
10.89%
8.01%
—
Gladia Solaria
incomplete
18.83%
25.30%
—
OpenAI GPT-4o mini Transcribe
incomplete
25.45%
19.27%
—
OpenAI GPT-4o Transcribe
incomplete
23.34%
14.91%
—
OpenAI GPT Realtime
incomplete
13.73%
5.39%
—
Best Others The best value in each column is filled blue.
Medication and disease
Table 6Medication and disease terms scored separately. We lead on the disease mean.medication / disease entity error rate · runs 2026-08-18 → 2026-09-02 · one row per model, most recent run
Medication and disease terms scored separately. We lead on the disease mean.
#
System model / snapshot
Medication mean
Primary care
Dictation
Respiratory
Disease mean
Primary care
Dictation
Respiratory
1
Speechmatics Enhanced Realtime
6.62%
8.92%
5.04%
5.90%
3.38%
7.35%
0.93%
1.86%
2
AssemblyAI Universal-3.5 Pro Realtime
8.70%
9.55%
7.31%
9.25%
2.52%
4.83%
1.17%
1.57%
3
Google Gemini 3.5 Transcribe Live
12.23%
19.19%
10.80%
6.70%
5.16%
11.76%
1.64%
2.09%
4
Deepgram Nova-3
17.76%
25.05%
20.53%
7.69%
8.83%
13.03%
11.21%
2.24%
5
Mistral Voxtral Mini Realtime
18.22%
18.47%
22.09%
14.10%
5.98%
6.93%
4.21%
6.79%
6
Azure Realtime STT
18.60%
23.57%
21.97%
10.26%
9.57%
10.92%
13.55%
4.25%
7
xAI Grok Streaming
19.50%
24.63%
26.05%
7.81%
7.74%
11.34%
6.31%
5.56%
8
Deepgram Flux EN
19.77%
29.30%
20.77%
9.23%
9.18%
11.76%
13.55%
2.24%
9
Speechmatics Standard Realtime
19.81%
26.98%
21.13%
11.31%
9.56%
10.53%
14.02%
4.13%
10
Reson8 Reson8
22.55%
27.74%
27.85%
12.05%
7.15%
8.80%
10.05%
2.61%
11
Smallest Pulse
26.83%
31.55%
41.42%
7.53%
12.04%
18.30%
10.05%
7.76%
AWS Transcribe Streaming
incomplete
15.07%
21.97%
—
incomplete
7.14%
5.84%
—
Cartesia Ink-Whisper
incomplete
33.55%
37.01%
—
incomplete
15.34%
19.57%
—
ElevenLabs Scribe v2 Realtime
incomplete
13.57%
11.64%
—
incomplete
8.17%
0.93%
—
Gladia Solaria
incomplete
23.81%
32.41%
—
incomplete
12.86%
11.45%
—
OpenAI GPT-4o mini Transcribe
incomplete
32.91%
21.49%
—
incomplete
18.07%
14.95%
—
OpenAI GPT-4o Transcribe
incomplete
27.18%
16.93%
—
incomplete
19.54%
10.98%
—
OpenAI GPT Realtime
incomplete
17.83%
6.72%
—
incomplete
9.66%
2.80%
—
Best Others The best value in each column is filled blue.
Public corpus of mock primary-care consultations: clinicians and actor-patients work through realistic appointments, recorded as telemedicine calls.
◆Medical dictation 93 recordings
Licensed clinician dictation, notes and findings spoken for the record, dense with medication and disease terms.
●Respiratory clinic recordings 272 recordings
Real clinical conversation from respiratory clinic encounters, with the accents, interruptions and room noise of a working clinic.
○public, open source ◆public, licensed ●private, ours
04English
General English: entity accuracy and non-speech handling first, then word error rate on short-form, accented and long-form sets.
Table 7Judge-scored entity error (%) on general English, by class. We lead on both overall scores. Per class, we lead on person names, product names, street addresses, organisations and URLs. A 0.00 cell means no data.LLM-judge-scored entity error rate (normalized and formatted) · runs 2026-08-18 → 2026-09-02 · one row per model, most recent run
Judge-scored entity error (%) on general English, by class. We lead on both overall scores. Per class, we lead on person names, product names, street addresses, organisations and URLs. A 0.00 cell means no data.
System
All classes
All, formatted
Persons
Products
Addresses
Orgs
Places
Phone
Email
URLs
Dates
Credit cards
Alphanum.
AssemblyAI Universal-3.5 Pro
19.2
27.8
19.9
20.9
24.0
18.4
21.3
14.3
20.8
14.2
21.8
13.7
8.0
ElevenLabs Scribe v2
21.2
29.2
24.8
25.5
32.6
23.2
19.1
n/d
19.9
14.9
22.0
n/d
9.2
Mistral Voxtral Mini
23.4
31.8
27.8
28.0
28.9
19.6
17.5
3.8
95.5
31.0
19.9
n/d
7.8
Speechmatics Enhanced
24.8
37.7
25.3
26.8
36.6
22.2
16.8
86.5
69.7
23.5
21.1
n/d
9.2
Reson8
25.3
38.4
37.0
27.2
31.4
23.6
18.5
8.0
16.3
26.0
22.9
4.2
9.3
OpenAI GPT-4o Transcribe
25.9
38.1
26.2
28.6
41.3
23.4
26.4
18.1
39.4
19.7
24.8
21.4
13.9
OpenAI GPT-4o mini Transcribe
26.6
38.6
28.7
29.4
41.3
24.2
27.3
8.9
45.7
20.4
23.2
8.3
14.4
OpenAI GPT Realtime
27.6
38.7
24.9
30.3
37.6
20.2
18.0
2.1
98.2
52.5
22.7
3.0
44.3
Gladia Solaria
28.0
41.8
32.5
30.4
38.4
24.3
22.4
4.2
95.9
32.3
24.1
7.7
11.3
Cartesia Ink-Whisper
29.8
43.3
37.7
30.5
39.4
27.6
26.3
12.7
53.4
25.4
24.8
13.7
12.8
Deepgram Nova-3
30.7
43.5
36.4
34.0
37.6
27.0
22.7
81.9
22.6
22.2
40.2
7.1
18.0
AWS Transcribe
30.7
44.1
46.2
33.6
26.8
28.2
23.9
0.4
79.6
28.7
22.9
8.3
7.7
Speechmatics Standard
32.3
44.6
34.0
29.9
41.3
26.8
20.6
85.7
92.3
41.0
22.1
n/d
36.9
xAI Grok Streaming
36.1
51.8
34.2
34.2
55.6
26.1
26.9
2.5
98.2
80.8
57.2
4.8
12.4
Smallest Pulse
40.7
58.4
53.2
37.2
71.8
32.4
30.2
3.8
23.1
25.3
24.8
5.4
61.4
Azure
50.2
97.5
46.0
42.0
89.6
36.8
30.5
2.1
98.2
72.9
66.3
0.6
100.0
Deepgram Flux EN
50.5
62.5
49.4
39.3
88.2
31.7
26.3
3.4
98.2
99.1
63.1
n/d
100.0
BestWorstCells are coloured by rank within each column (solid blue is best, white is worst).
Figure 7Non-speech response rate: how often a system produces text for a clip with no speech. Lower is better. The axis is clipped at 5%; the notched bar exceeds it.non-speech response rate (share of pure non-speech clips that produce any text) · runs 2026-08-18 → 2026-09-02 · one row per model, most recent run
Word error rate — short-form
Table 8Short-form English WER. We lead on the mean and Common Voice. Systems without a Common Voice run have no mean.normalized word error rate · runs 2026-08-18 → 2026-09-02 · one row per model, most recent run
Short-form English WER. We lead on the mean and Common Voice. Systems without a Common Voice run have no mean.
#
System model / snapshot
Mean 3 sets
Common Voice
LibriSpeech clean
LibriSpeech other
1
AssemblyAI Universal-3.5 Pro Realtime
3.87%
6.44%
1.88%
3.28%
2
Speechmatics Enhanced Realtime
4.69%
6.78%
2.31%
4.99%
3
Smallest Pulse
5.41%
9.34%
2.19%
4.69%
4
Azure Realtime STT
5.54%
8.68%
2.44%
5.49%
5
xAI Grok Streaming
5.81%
11.04%
2.04%
4.36%
6
Speechmatics Standard Realtime
6.41%
9.43%
3.18%
6.61%
7
Deepgram Flux EN
6.53%
7.68%
3.56%
8.35%
8
Mistral Voxtral Mini Realtime
6.61%
12.23%
2.10%
5.49%
9
Deepgram Nova-3
7.46%
12.38%
3.28%
6.72%
AWS Transcribe Streaming
incomplete
—
1.90%
3.77%
Cartesia Ink-Whisper
incomplete
—
3.19%
6.48%
ElevenLabs Scribe v2 Realtime
incomplete
—
1.96%
4.43%
Gladia Solaria
incomplete
—
2.73%
5.42%
OpenAI GPT-4o mini Transcribe
incomplete
—
2.53%
7.97%
OpenAI GPT-4o Transcribe
incomplete
—
2.20%
7.15%
OpenAI GPT Realtime
incomplete
—
2.59%
7.44%
Reson8 Reson8
incomplete
—
1.44%
2.68%
Best Others The best value in each column is filled blue.
Listen to the test data
Two read sentences from the short-form sets above, human reference beside each transcript. Words that differ from it are marked.
1 of 2
Short-form English17 s · Common Voice (CC0)
0:00 / 0:00
A read sentence. The other system appends a sentence and a name that are not in the audio.
Reference
You can do a lot just using your voice, but there are still a few times you'll find yourself reaching for a mouse.
AssemblyAI
You can do a lot just using your voice. But there are still a few times you'll find yourself reaching for a mouse.
Speechmatics Enhanced
I know you can do a lot just using your voice, but there are still a few times you'll find yourself reaching for a mouse. And I'm Larry Hanson. We're here in this beautiful hall.
Word error rate — accented English
Table 9Accented English WER. We lead on British-accented English and Indian-accented English.normalized word error rate · runs 2026-08-18 → 2026-09-02 · one row per model, most recent run
Accented English WER. We lead on British-accented English and Indian-accented English.
#
System model / snapshot
Mean 4 sets
British
French-acc.
Spanish-acc.
Indian-acc.
1
Azure Realtime STT
6.09%
6.25%
4.44%
5.52%
8.15%
2
AssemblyAI Universal-3.5 Pro Realtime
7.47%
5.67%
8.67%
9.60%
5.94%
3
Reson8 Reson8
9.25%
5.76%
11.89%
12.60%
6.73%
4
xAI Grok Streaming
9.85%
7.33%
12.21%
12.88%
6.99%
5
ElevenLabs Scribe v2 Realtime
10.10%
6.92%
12.12%
13.29%
8.07%
6
OpenAI GPT-4o Transcribe
10.30%
7.27%
11.83%
13.89%
8.22%
7
Speechmatics Enhanced Realtime
10.44%
6.70%
13.26%
14.80%
6.98%
8
AWS Transcribe Streaming
11.10%
7.02%
15.49%
15.47%
6.42%
9
Cartesia Ink-Whisper
11.23%
8.16%
12.92%
14.57%
9.27%
10
OpenAI GPT-4o mini Transcribe
11.36%
7.88%
13.31%
15.35%
8.88%
11
OpenAI GPT Realtime
11.78%
8.31%
14.83%
15.78%
8.20%
12
Smallest Pulse
11.85%
7.70%
15.15%
16.53%
8.03%
13
Deepgram Nova-3
11.96%
9.14%
13.92%
15.19%
9.57%
14
Deepgram Flux EN
12.55%
9.61%
14.83%
16.12%
9.64%
15
Speechmatics Standard Realtime
12.58%
8.43%
16.29%
17.46%
8.14%
16
Mistral Voxtral Mini Realtime
12.63%
7.39%
14.19%
15.87%
13.06%
Gladia Solaria
incomplete
8.67%
12.60%
13.85%
—
Best Others The best value in each column is filled blue.
Word error rate — long-form
Table 10Long-form English WER. TED-LIUM references omit sponsor read-outs present in the audio.normalized word error rate · runs 2026-08-18 → 2026-09-02 · one row per model, most recent run
Long-form English WER. TED-LIUM references omit sponsor read-outs present in the audio.
#
System model / snapshot
Mean 3 sets
TED-LIUM
Rev16
CALLHOME per channel
1
Reson8 Reson8
6.86%
4.63%
6.87%
9.09%
2
Speechmatics Enhanced Realtime
7.74%
6.29%
7.34%
9.59%
3
AWS Transcribe Streaming
8.03%
6.45%
7.81%
9.82%
4
Speechmatics Standard Realtime
8.55%
6.83%
8.06%
10.77%
5
Azure Realtime STT
8.73%
7.16%
8.17%
10.87%
6
Smallest Pulse
8.99%
5.23%
7.62%
14.13%
7
xAI Grok Streaming
9.04%
6.30%
8.00%
12.82%
8
Deepgram Flux EN
9.07%
7.01%
8.08%
12.12%
9
Deepgram Nova-3
9.34%
7.43%
9.40%
11.19%
10
AssemblyAI Universal-3.5 Pro Realtime
9.53%
8.63%
8.36%
11.61%
11
Cartesia Ink-Whisper
11.06%
7.13%
10.42%
15.62%
12
Gladia Solaria
11.68%
7.75%
9.04%
18.25%
13
ElevenLabs Scribe v2 Realtime
19.95%
5.96%
8.38%
45.51%
14
Mistral Voxtral Mini Realtime
22.73%
6.96%
8.33%
52.91%
OpenAI GPT-4o mini Transcribe
incomplete
7.28%
21.26%
—
OpenAI GPT-4o Transcribe
incomplete
9.88%
28.51%
—
OpenAI GPT Realtime
incomplete
5.70%
6.02%
—
Best Others The best value in each column is filled blue.
Data
●Judge-scored entities
The short-form English audio below. An LLM judge grades each entity (person, product, organisation, place, address, phone, email, date, ID, URL) on whether the transcript names the same real-world thing, with and without formatting.
●Non-speech 2,950 clips
Hold music, line noise and silence, with no words at all. The score is how often a model emits text anyway.
○Common Voice 16,029 clips
Crowdsourced read sentences from Mozilla’s public Common Voice, recorded by volunteers on their own phones and laptops, so accents, microphones and rooms vary widely.
○LibriSpeech test-clean 2,620 utterances
Read English audiobook speech from the public LibriSpeech corpus, recorded close-mic by LibriVox volunteers. The "clean" split is the easier half: clear speakers, little noise.
○LibriSpeech test-other 2,939 utterances
The harder half of the same LibriSpeech set: speakers and recordings that automatic transcription finds more difficult, still read prose.
◆Accented English 9,315 files
Four licensed sets of prompted British-, French-, Spanish- and Indian-accented English: accuracy outside the General American accent most models are tuned on.
○TED-LIUM 11 full-length talks
Full TED talks from the public TED-LIUM corpus: one prepared speaker on stage, recorded end to end, many minutes of continuous speech per file.
○Rev16 9 full-length recordings
Podcast-style recordings from the public Rev16 set: spontaneous multi-speaker studio conversation, episode length.
◆CALLHOME 71 scored channels
Public CALLHOME telephone calls between friends and family: unscripted, fast, overlapping conversation over a real phone line. Each speaker's channel is scored on its own.
○public, open source ◆public, licensed ●private, ours
05Multilingual
Nine languages on FLEURS and Common Voice; every system is compared on the same nine.
FLEURS
Table 11WER on FLEURS (CER for languages marked *). The nine-language mean is shown only for complete runs.normalized word error rate (character error rate for Japanese, Chinese, Korean) · runs 2026-08-25 → 2026-09-03 · one row per model, most recent run
WER on FLEURS (CER for languages marked *). The nine-language mean is shown only for complete runs.
System
Mean 9 lang., excl. ko
Spanish
French
German
Italian
Portuguese
Dutch
Hindi
Japanese*
Chinese*
Korean*
AssemblyAI Universal-3.5 Pro
7.2
4.3
5.2
5.3
4.5
5.7
6.7
12.5
6.7
13.9
n/s
AWS Transcribe
7.4
5.3
7.6
8.0
4.6
8.1
7.1
6.3
7.0
12.3
4.9
ElevenLabs Scribe v2
7.7
4.6
6.9
5.8
4.3
6.4
7.0
15.8
8.6
10.1
4.6
Mistral Voxtral Mini
9.7
4.8
9.6
7.8
5.5
6.9
9.5
16.8
9.6
16.5
5.8
Azure
9.8
7.7
11.4
9.1
8.3
10.5
12.9
8.6
8.6
11.5
9.4
OpenAI GPT Realtime
10.4
5.7
11.8
11.2
5.9
8.8
10.1
10.4
13.4
16.3
7.1
OpenAI GPT-4o Transcribe
11.2
4.3
10.6
13.1
5.5
8.5
13.0
12.4
16.1
17.4
6.7
Smallest Pulse
12.0
8.9
12.7
11.6
7.0
11.3
13.1
10.7
12.7
20.1
7.4
Speechmatics Enhanced
12.4
4.7
6.5
6.8
5.4
7.3
8.5
7.0
31.3
33.8
5.4
Speechmatics Standard
13.0
5.4
8.1
8.2
5.7
7.3
9.5
7.0
31.4
34.7
5.2
OpenAI GPT-4o mini Transcribe
13.3
5.5
13.2
14.9
6.9
10.9
16.8
16.8
15.4
19.3
7.6
Gladia Solaria
14.1
7.3
12.7
10.9
7.5
8.7
11.6
23.9
23.6
21.1
6.6
Cartesia Ink-Whisper
16.8
9.0
16.9
12.7
9.7
10.7
17.4
27.5
20.4
26.3
10.9
Soniox Realtime
incomplete
5.0
7.8
6.8
—
—
7.6
—
—
—
4.4
Deepgram Nova-3
incomplete
8.9
14.0
13.3
10.0
13.3
14.6
22.3
18.8
n/s
n/s
Deepgram Flux
incomplete
8.1
13.9
10.4
9.2
12.7
16.4
20.5
14.4
n/s
n/s
BestWorstCells are coloured by rank within each column (solid blue is best, white is worst).
Common Voice
Table 12WER on Common Voice (CER for languages marked *).normalized word error rate (character error rate for Japanese, Chinese, Korean) · runs 2026-08-25 → 2026-09-03 · one row per model, most recent run
WER on Common Voice (CER for languages marked *).
System
Mean 9 lang., excl. ko
Spanish
French
German
Italian
Portuguese
Dutch
Hindi
Japanese*
Chinese*
Korean*
AWS Transcribe
6.6
4.1
8.8
5.8
4.2
4.6
2.7
5.1
15.9
8.5
7.5
Speechmatics Enhanced
8.7
2.5
7.2
3.7
4.7
5.2
2.4
3.9
23.2
25.1
7.2
AssemblyAI Universal-3.5 Pro
8.7
4.4
8.4
5.0
4.5
8.1
3.7
10.6
19.9
13.7
n/s
Speechmatics Standard
10.1
4.2
10.2
6.2
5.3
6.4
3.0
4.0
25.4
25.8
8.3
OpenAI GPT Realtime
10.9
6.9
12.6
9.3
7.8
9.4
6.8
8.4
20.2
16.4
10.5
ElevenLabs Scribe v2
11.0
5.8
10.2
7.1
6.1
9.7
5.5
21.6
20.7
12.1
11.0
Smallest Pulse
12.8
8.6
12.3
11.3
6.6
21.6
9.2
9.7
17.5
18.0
10.7
OpenAI GPT-4o Transcribe
13.2
8.9
15.5
10.3
9.2
13.2
7.8
10.0
22.0
21.4
12.1
Mistral Voxtral Mini
13.7
5.7
11.5
8.8
7.4
9.9
8.5
18.1
22.8
30.5
14.3
OpenAI GPT-4o mini Transcribe
15.4
9.8
16.8
11.7
10.6
14.3
11.9
21.5
22.3
19.3
12.2
Cartesia Ink-Whisper
17.6
10.4
17.7
12.4
11.1
11.5
9.7
27.9
28.3
29.2
13.8
Soniox Realtime
incomplete
6.4
12.2
7.6
7.2
8.6
4.6
—
19.2
14.1
8.7
Azure
incomplete
6.8
11.0
6.1
—
6.8
4.6
4.9
19.8
9.9
10.4
Gladia Solaria
incomplete
8.1
14.4
10.5
9.8
9.5
—
22.0
26.2
29.0
9.6
Deepgram Nova-3
incomplete
8.4
13.8
12.8
10.8
12.7
9.3
25.2
29.7
n/s
n/s
Deepgram Flux
incomplete
9.1
14.2
11.0
13.9
12.5
11.2
23.4
29.9
n/s
n/s
BestWorstCells are coloured by rank within each column (solid blue is best, white is worst).
Data
○FLEURS
Public corpus of parallel sentences read by native speakers in quiet settings, one test set per language. Shown: Spanish, French, German, Italian, Portuguese, Dutch, Hindi, Japanese and Chinese.
○Common Voice
Mozilla’s public Common Voice, per language: volunteers reading sentences on their own devices, so accents and microphones vary. Japanese and Chinese are scored by character error rate.
○public, open source ◆public, licensed ●private, ours
06Code-switching
Speech that switches between English and a second language within one utterance: five language pairs plus two Spanish–English conversational corpora (Bangor Miami).
Table 13Code-switching WER. We lead on the five-pair mean. We lead in 4 of 5 pairs.normalized word error rate · runs 2026-08-18 → 2026-09-06 · one row per model, most recent run · pair value = mean of 4 sub-sets
Code-switching WER. We lead on the five-pair mean. We lead in 4 of 5 pairs.
#
System model / snapshot
Mean 5 pairs
EN↔ES
EN↔DE
EN↔FR
EN↔IT
EN↔PT
Bangor Miami
Miami 10h
1
AssemblyAI Universal-3.5 Pro Realtime
7.82%
6.80%
7.84%
8.54%
7.21%
8.71%
31.65%
27.31%
2
AWS Transcribe Streaming
8.30%
7.17%
8.12%
9.33%
7.05%
9.81%
31.65%
27.96%
3
Soniox Realtime
10.77%
9.65%
10.82%
12.11%
10.09%
11.16%
26.84%
21.97%
4
ElevenLabs Scribe v2 Realtime
12.37%
11.06%
12.29%
13.16%
11.53%
13.79%
27.77%
27.09%
5
Smallest Pulse
13.13%
11.40%
13.14%
13.61%
10.79%
16.73%
43.62%
47.17%
6
Azure Realtime STT
13.75%
13.02%
13.33%
14.86%
12.52%
15.03%
34.83%
34.17%
7
Gladia Solaria
16.28%
14.36%
18.10%
18.73%
14.28%
15.93%
35.50%
34.49%
8
Deepgram Nova-3 Multi
16.47%
13.36%
18.01%
18.52%
15.17%
17.29%
31.04%
29.91%
9
Mistral Voxtral Mini Realtime
17.67%
14.40%
19.11%
20.79%
16.61%
17.43%
34.76%
34.35%
10
OpenAI GPT-4o mini Transcribe
32.78%
32.84%
33.06%
34.22%
31.73%
32.07%
48.24%
45.01%
11
OpenAI GPT-4o Transcribe
34.38%
34.26%
34.70%
35.78%
33.17%
34.00%
61.76%
61.27%
12
OpenAI GPT Realtime
36.65%
35.85%
35.46%
37.87%
36.30%
37.77%
55.84%
53.31%
13
Speechmatics Enhanced Realtime
38.11%
33.18%
33.78%
35.98%
45.39%
42.23%
43.51%
36.04%
14
Speechmatics Standard Realtime
42.70%
36.77%
44.12%
42.99%
45.11%
44.49%
48.14%
40.95%
15
Cartesia Ink-Whisper EN
50.69%
47.71%
51.50%
53.10%
53.17%
47.95%
52.94%
43.72%
16
Deepgram Nova-3 EN
55.23%
57.02%
54.61%
58.15%
56.24%
50.15%
52.21%
43.85%
Best Others The best value in each column is filled blue.
Figure 8Five-pair mean WER.normalized word error rate · runs 2026-08-18 → 2026-09-06 · one row per model, most recent run
Audio that alternates between English and the second language mid-stream, built from public read-speech corpora in short, long and split-long clips so the switch falls at different points.
○Bangor Miami 56 recordings
Public corpus of spontaneous Spanish-English conversation among Miami residents who switch language mid-sentence. The hardest audio in this suite.
◆Miami, 10-hour selection 21 recordings
A ten-hour selection of the Miami Spanish-English conversation, with references produced independently by a third-party transcription service.
○public, open source ◆public, licensed ●private, ours
07Diarization
Speaker attribution on CALLHOME, AMI, DiPCo and NOTSOFAR, for systems that offer streaming diarization. Diarization error rate is the share of audio time given to the wrong speaker, missed, or wrongly marked as speech.
Table 14Diarization error rate and cpWER. We lead on the DER mean and the cpWER mean.diarization error rate (DER) and cpWER · runs 2026-08-18 → 2026-09-01 · one row per model, most recent run
Diarization error rate and cpWER. We lead on the DER mean and the cpWER mean.
#
System model / snapshot
DER mean 4 sets
CALLHOME
AMI
DiPCo
NOTSOFAR
cpWER mean
CALLHOME
AMI
DiPCo
NOTSOFAR
1
AssemblyAI Universal-3.5 Pro Realtime
17.62%
13.82%
19.61%
25.57%
11.49%
32.65%
19.71%
30.09%
38.99%
41.81%
2
Azure Realtime STT
20.43%
20.21%
18.00%
24.30%
19.19%
40.96%
34.68%
33.63%
41.45%
54.06%
3
Speechmatics Enhanced Realtime
22.16%
14.58%
19.96%
28.37%
25.74%
37.14%
24.69%
23.01%
42.59%
58.25%
4
xAI Grok Streaming
25.68%
15.59%
31.90%
35.28%
19.96%
41.81%
25.77%
37.95%
44.72%
58.82%
5
Speechmatics Standard Realtime
25.71%
25.43%
21.27%
24.68%
31.47%
48.86%
47.84%
33.66%
43.43%
70.49%
6
AWS Transcribe Streaming
27.48%
27.89%
26.16%
32.09%
23.79%
46.07%
41.66%
38.72%
47.91%
55.98%
7
Soniox Realtime
31.46%
20.29%
44.53%
27.56%
33.45%
42.52%
18.30%
54.27%
39.45%
58.08%
8
Deepgram Nova-3
34.01%
30.46%
33.86%
45.35%
26.37%
56.09%
52.17%
43.99%
69.88%
58.30%
9
Smallest Pulse
41.35%
27.16%
59.41%
50.98%
27.87%
53.34%
31.65%
65.89%
62.56%
53.26%
Best Others The best value in each column is filled blue.
Data
◆CALLHOME 40 calls
Public CALLHOME telephone calls between friends and family: unscripted, fast, heavily overlapping conversation over a real phone line.
○AMI 24 meetings
Public AMI Meeting Corpus: small groups of colleagues in real design and project meetings, in instrumented rooms.
○NOTSOFAR 129 meetings
Public NOTSOFAR corpus: real multi-party meetings on a single distant microphone in ordinary meeting rooms.
○DiPCo 10 sessions
Public Dinner Party Corpus: groups talking over dinner on far-field microphones, constant overlap and background clatter.
○public, open source ◆public, licensed ●private, ours
08Turn detection (EoT-Bench)
End-of-turn detection on LiveKit’s public EoT-Bench (English), separate from our voice-agent set. Our production model is included.
Table 15EoT-Bench English: end-of-turn F1, precision, recall and endpoint latency. We lead on recall.EoT-Bench (LiveKit), English · runs 2026-08-31 → 2026-09-02 · fractions ×100
EoT-Bench English: end-of-turn F1, precision, recall and endpoint latency. We lead on recall.
#
System model / snapshot
F1 %
Precision %
Recall %
P50 ms
P90 ms
Mean cutoff ms
1
OpenAI GPT Realtime
99.3
99.3
99.3
1574
1861
7
2
Mistral Voxtral Mini Realtime
97.8
97.8
97.8
1564
1867
6
3
Google Gemini 3.5 Transcribe Live
95.6
94.3
98.8
927
1143
9
4
ElevenLabs Scribe v2 Realtime
95.1
93.9
98.0
1360
1617
6
5
Cartesia Ink 2 (turns)
93.9
93.3
95.3
538
910
47
6
AssemblyAI Universal-3.5 Pro Realtime
90.3
86.8
99.8
438
1359
16
7
Azure Realtime STT
88.5
85.5
96.5
536
721
31
8
AWS Transcribe Streaming
86.7
82.4
98.8
499
684
20
9
Soniox Realtime
86.5
84.8
90.3
274
966
95
10
Deepgram Flux EN
84.8
83.1
88.5
270
631
139
11
Deepgram Nova-3
77.3
75.0
83.5
323
627
65
12
Speechmatics Enhanced Realtime
76.8
74.2
83.5
1046
1343
15
13
OpenAI GPT-4o mini Transcribe
71.0
64.1
95.5
552
783
63
14
Smallest Pulse
70.9
63.7
95.0
586
929
64
15
OpenAI GPT-4o Transcribe
70.8
63.7
96.3
569
780
65
16
Gladia Solaria
65.7
58.1
92.5
648
966
72
17
Speechmatics Standard Realtime
45.3
43.7
49.0
1004
1341
14
Best Others The best value in each column is filled blue.
Data
○EoT-Bench (LiveKit), English 400 clips
The English test set of LiveKit's public EoT-Bench: conversational turn-taking clips with human-marked turn boundaries.
○public, open source ◆public, licensed ●private, ours
09Independent evaluations
Third-party results on their own corpora and pipelines, reproduced for comparison.
Artificial Analysis — Speech to Text (Streaming)
AA-WER v2.2 (May 2026): 50% AA-AgentTalk, 25% VoxPopuli, 25% Earnings22. End of speech is detected with SileroVAD; latency runs from there to the final transcript. Price is per 1,000 minutes of audio, normalized across billing models. The leaderboard publishes no “last updated” date.
Table 1630 of the 31 systems on the Artificial Analysis streaming leaderboard, at published ranks; our earlier-generation model is omitted.source = https://artificialanalysis.ai/speech-to-text/streaming · read 2026-09-09 · price shown for every system the source publishes one for
30 of the 31 systems on the Artificial Analysis streaming leaderboard, at published ranks; our earlier-generation model is omitted.
#
System
AA-WER index %
t final s
Price $/1k min
1
Muse Voice Transcribe (Meta)
3.06%
0.163 s
3.0
2
Cartesia Ink-2 (semantic)
3.36%
0.431 s
4.0
3
ElevenLabs Scribe v2 Realtime
3.59%
0.141 s
6.5
4
Qwen3 ASR Flash Realtime
3.73%
0.476 s
5.4
5
GPT Live Transcribe
3.92%
0.812 s
17.0
6
Grok STT Streaming
3.93%
0.373 s
3.3
7
Gemini 3.5 Transcribe Live
4.00%
0.395 s
9.0
8
AssemblyAI U3.5 Realtime Pro — min latency
4.02%
0.191 s
7.5
9
Cartesia Ink-2 (external endpoint)
4.02%
0.067 s
4.0
10
AssemblyAI U3.5 Realtime Pro — max accuracy
4.03%
0.202 s
7.5
12
Inworld STT 1
4.18%
0.071 s
1.4
13
Soniox v4
4.49%
0.062 s
2.0
14
Soniox v5
4.50%
0.054 s
2.0
15
Google Chirp 3 Streaming
4.80%
1.276 s
16.0
16
OpenAI GPT Realtime (Whisper)
4.89%
0.688 s
17.0
17
Smallest Pulse
4.98%
0.132 s
4.0
18
Voxtral Mini Realtime
5.24%
0.682 s
6.0
19
Azure Speech
5.25%
0.625 s
16.7
20
Nemotron 3 ASR (1120 ms)
5.36%
0.418 s
—
21
Amazon Transcribe
5.79%
0.620 s
24.0
22
Nemotron 3 ASR (560 ms)
5.96%
0.252 s
—
23
Deepgram Nova-3 Realtime
6.59%
0.066 s
4.8
24
Gradium
6.61%
0.435 s
12.0
25
Nemotron 3 ASR (80 ms, Together)
7.03%
0.070 s
—
26
Deepgram Flux
7.39%
0.021 s
6.5
27
Nemotron 3 ASR (160 ms)
7.61%
0.098 s
—
28
Gladia Solaria
7.82%
1.523 s
12.5
29
Speechmatics Realtime Enhanced
8.05%
0.317 s
17.5
30
Nemotron 3 ASR (80 ms)
8.38%
0.073 s
—
31
Rev.ai
10.14%
0.697 s
3.3
Best Others The best value in each column is filled blue.
Table 17The three sets in the Artificial Analysis index. The source ranks our two configurations separately; both are shown beside each set’s leading system.source = https://artificialanalysis.ai/speech-to-text/streaming · read 2026-09-09
The three sets in the Artificial Analysis index. The source ranks our two configurations separately; both are shown beside each set’s leading system.
Set
Our configuration
WER
Rank
Leading system on the set
Its WER
VoxPopuli
U3.5 RT Pro — min latency
2.06%
3 of 31
Gemini 3.5 Transcribe Live
1.76%
VoxPopuli
U3.5 RT Pro — max accuracy
2.09%
4 of 31
Gemini 3.5 Transcribe Live
1.76%
AA-AgentTalk
U3.5 RT Pro — min latency
3.53%
9 of 31
Qwen3 ASR Flash Realtime
2.49%
AA-AgentTalk
U3.5 RT Pro — max accuracy
3.48%
7 of 31
Qwen3 ASR Flash Realtime
2.49%
Earnings22
U3.5 RT Pro — min latency
6.96%
11 of 31
Muse Voice Transcribe
4.31%
Earnings22
U3.5 RT Pro — max accuracy
7.08%
12 of 31
Muse Voice Transcribe
4.31%
1 ElevenLabs Scribe v2 Realtime · 0.141 s · 3.59%
2 Gemini 3.5 Transcribe Live · 0.395 s · 4.00%
3 Google Chirp 3 Streaming · 1.276 s · 4.80%
4 OpenAI GPT Realtime (Whisper) · 0.688 s · 4.89%
5 Nemotron 3 ASR (560 ms) · 0.252 s · 5.96%
Figure 9AA-WER against time to final transcript. The dashed line is the accuracy–latency frontier.source = https://artificialanalysis.ai/speech-to-text/streaming
Coval — Live STT Leaderboard
Rolling 24-hour window (the source’s default view), read 2026-09-09 with 28 systems ranked. WER is pooled over seven conditions (clean, accent, clipping, far-field, noise gap, phone codec, reverb), 44–45 clips each and 176 for clean. The stt-v3 references are model-generated, not human-verified. Coval sells voice-agent evaluation tooling.
Table 18Selected rows of 28. We rank 1st of 28 on pooled WER.source = https://benchmarks.coval.ai/stt · read 2026-09-09 · rolling 24 h window
Selected rows of 28. We rank 1st of 28 on pooled WER.
System model / snapshot
WER pooled %
AssemblyAI universal-3.5-pro rank 1 of 28
2.91%
Baseten qwen3-asr-1.7b rank 2 of 28
3.19%
Reson8 realtime rank 3 of 28
3.34%
Google Chirp 3 rank 4 of 28
4.15%
Inworld rank 5 of 28
4.18%
ElevenLabs Scribe v2 Realtime rank 16 of 28
5.69%
Soniox stt-rt-v5 rank 19 of 28
6.21%
Deepgram Flux (general-en) rank 21 of 28
6.67%
Best Others The best value in each column is filled blue.
Figure 10Our WER by recording condition. The dashed line is the pooled value from Table 18. The axis is clipped at 4%; reverb exceeds it.source = https://benchmarks.coval.ai/stt
Pipecat / Daily — STT benchmark
1,000 English samples. Semantic WER is pooled over all reference words; time to final is the median and P95 from end of speech to the final segment. Read from the repository README (last changed 2026-08-27). Vendors without a model name in the source are listed by vendor only; previous-generation AssemblyAI models are not shown.
Table 19Shown as published; our row is the production model named in the source.source = https://github.com/pipecat-ai/stt-benchmark · read 2026-09-09
Shown as published; our row is the production model named in the source.
#
System
Semantic WER %
TTF median ms
TTF p95 ms
1
Speechmatics
1.07%
495
676
2
Azure
1.18%
1016
1345
3
AssemblyAI universal-3.5-pro
1.22%
282
354
4
Cartesia ink-2
1.25%
299
328
5
Soniox stt-rt-v5
1.27%
260
305
6
Soniox stt-rt-v4
1.29%
249
281
7
Deepgram nova-3-general
1.62%
247
298
8
AWS Transcribe
1.75%
1136
1527
9
NVIDIA Nemotron 3.0 ASR (en)
1.95%
221
238
10
Google gemini-3.5-transcribe-live
2.24%
458
532
11
Smallest AI pulse
2.37%
398
533
12
OpenAI gpt-realtime-whisper
2.73%
740
878
13
Google latest-long
2.85%
878
1155
14
OpenAI gpt-4o-transcribe
3.06%
637
965
15
ElevenLabs scribe_v2_realtime
3.12%
281
348
16
Gradium
3.71%
570
596
17
Cartesia ink-whisper
4.36%
266
364
18
NVIDIA Nemotron 3.5 ASR (multilingual)
4.58%
236
253
19
Mistral voxtral-mini-transcribe-realtime-2602
4.97%
525
973
Best Others The best value in each column is filled blue.
10How we measure
How the results were produced: the harness, the metrics, why public leaderboards can rank the same systems differently, and the limits of this report.
How we run every system
Protocol. Audio is streamed in real time to each system’s public API, default settings, language fixed. One normalizer is applied before scoring. Each streaming model a vendor ships has its own row.
Complete runs only. An average is shown only when the system was scored on every dataset in it; partial rows keep their cells but are not ranked.
n/d, n/s. A 0.00 error rate means no data. WER or CER of 60% or more means the language is unsupported; the cell is excluded from means.
Excluded. Unreleased models, and results on a superseded version of the voice-agent test set.
Systems
Every system with a public streaming API and a current run; a dot marks each suite it was scored on.
Table 20Systems evaluated and the suites each has a current run on.6 suites · runs 2026-08-27 → 2026-09-08
Systems evaluated and the suites each has a current run on.
Vendor
Model
English / CI
Voice agents
Code-switch.
Diarization
Multilingual
Turn detect.
Runs
AssemblyAI
Universal-3.5 Pro Realtime
●
●
●
●
●
●
2026-08-18 → 2026-09-02
AWS
Transcribe Streaming
●
●
●
●
●
●
2026-08-19 → 2026-08-31
Azure
Azure Realtime STT
●
·
●
●
●
●
2026-08-19 → 2026-08-31
Cartesia
Ink 2 (turns)
·
·
·
·
·
●
2026-08-31
Cartesia
Ink-Whisper
●
●
·
·
●
·
2026-08-19 → 2026-08-30
Cartesia
Ink-Whisper EN
·
·
●
·
·
·
2026-08-20
Deepgram
Flux
·
·
·
·
●
·
2026-08-25
Deepgram
Flux EN
●
●
·
·
·
●
2026-08-19 → 2026-08-31
Deepgram
Flux Multi
·
●
·
·
·
·
2026-08-27
Deepgram
Nova-3
●
●
·
●
●
●
2026-08-18 → 2026-08-31
Deepgram
Nova-3 EN
·
·
●
·
·
·
2026-08-20
Deepgram
Nova-3 Multi
·
·
●
·
·
·
2026-08-20
ElevenLabs
Scribe v2 Realtime
●
●
●
·
●
●
2026-08-19 → 2026-09-01
Gladia
Solaria
●
·
●
·
●
●
2026-08-19 → 2026-08-31
Google
Gemini 3.5 Transcribe Live
●
●
●
·
●
●
2026-08-31 → 2026-09-06
Mistral
Voxtral Mini Realtime
●
●
●
·
●
●
2026-08-19 → 2026-08-31
OpenAI
GPT Realtime
●
●
●
·
●
●
2026-08-19 → 2026-08-31
OpenAI
GPT-4o Transcribe
●
●
●
·
●
●
2026-08-19 → 2026-08-31
OpenAI
GPT-4o mini Transcribe
●
●
●
·
●
●
2026-08-19 → 2026-08-31
Reson8
Reson8
●
·
·
·
·
·
2026-09-01
Smallest
Pulse
●
●
●
●
●
●
2026-08-28 → 2026-08-31
Soniox
Soniox Realtime
·
●
●
●
●
●
2026-08-20 → 2026-08-31
Speechmatics
Enhanced Realtime
●
●
●
●
●
●
2026-08-19 → 2026-08-31
Speechmatics
Standard Realtime
●
●
●
●
●
●
2026-08-19 → 2026-08-31
xAI
Grok Streaming
●
●
·
●
·
·
2026-08-19 → 2026-09-08
Test data
The audio each benchmark uses and the metric it scores. Each results section also lists its own datasets.
Table 21Test data by benchmark. Files are scored recordings or clips; multilingual, code-switching and judge-scored sets are scored per language or label, so no single count applies. Public corpora are named, licensed sets described.
Test data by benchmark. Files are scored recordings or clips; multilingual, code-switching and judge-scored sets are scored per language or label, so no single count applies. Public corpora are named, licensed sets described.
Suite
Audio
Files
Scored on
Voice agents
Scripted caller scenarios voiced by real speakers, recorded in three noise environments, three utterance-length bands and 12 speaker groups
Short-form English audio, with entities graded per label by an LLM judge
—
Entity error (judge)
Non-speech
Hold music, line noise and silence with no speech
2,950
Response rate
Multilingual
FLEURS and Common Voice test sets in nine languages
—
WER (character error rate for Japanese, Chinese)
Code-switching
Five English↔X pairs derived from read speech, plus Bangor Miami bilingual conversation
—
WER
Diarization
CALLHOME calls; AMI and NOTSOFAR meetings; DiPCo dinner parties
203
DER, cpWER
Turn detection
EoT-Bench (LiveKit), English
400
End-of-turn F1, endpoint latency
Each test set is marked by how its audio can be obtained: ○ public and open source; ◆ public but licensed or purchasable; ● private, recorded or annotated by us.
How we score
Each metric and how it is computed. The same normalizer and scoring code apply to every system.
WER Word error rate after normalization; lower is better. Character error rate for Japanese, Chinese and Korean.
NEER Normalized entity error rate: the share of entities (names, addresses, numbers, emails, dates, codes, medical and technical terms) transcribed incorrectly.
MER Medical entity error rate over medication and disease terms.
Judge-scored entity error Entity error rate as graded by an LLM judge.
End-of-turn F1 / P / R Whether the endpoint fires exactly once per real turn. Latency is the time (ms) from true end of speech to the endpoint event.
DER / cpWER Diarization error rate, and concatenated minimum-permutation WER.
Non-speech response rate Share of non-speech clips for which the system produces any text.
Streaming harness. Audio is chunked and streamed in real time to each vendor’s endpoint through its documented SDK or WebSocket API, default settings, language fixed. Final transcripts are scored, not partials.
Normalization. One normalizer (case, punctuation, numbers, abbreviations) is applied to references and to every system’s output. Entity metrics are computed on normalized text and, where reported, on formatted text.
Entity error. Entity spans are located in the reference; an entity is an error if any token of its span is wrong in the aligned transcript. The judge-scored variant asks an LLM whether the transcript names the same real-world entity.
Endpointing. An endpoint is a true positive if it falls inside the tolerance window after the reference end of turn. Latency is its offset from that end of turn.
Selection rule. Every row is the model’s most recent production run. The overview grid keeps each vendor’s best model per column and names it on hover.
Provenance. Each figure carries its metric and run-date span. Public test sets are named; licensed sets are described.
Why public leaderboards and ours differ
We build for voice agents, clinical documentation, meetings, telephony, and multilingual and code-switched speech, so the test sets cover that range of speakers, accents, languages and acoustic conditions and score what matters in each. Public leaderboards usually ask a narrower question, typically one word-error-rate figure on a few corpora, so a model can rank differently on the two.
Table 22What each benchmark measures: this report and the public leaderboards reproduced in §9.
What each benchmark measures: this report and the public leaderboards reproduced in §9.
Benchmark
Audio
Scores
Run by
References
This report
Voice-agent calls in three noise environments, clinical consultations and dictation, accented and long-form English, nine languages, code-switched speech, meetings and telephone calls, non-speech audio
Entity error (by class), WER, DER, end-of-turn F1 and latency, non-speech response rate
One WER index over the three sets; time to final transcript
Artificial Analysis
Published with the sets
Coval
Short clips under seven recording conditions (clean, accent, far-field, phone codec, reverb…), rolling 24-hour window
Pooled WER; time to first token; time to final segment
Coval
Model-generated
Pipecat / Daily
1,000 English samples
Semantic WER; median and P95 time to final
Daily
Published with the set
EoT-Bench (LiveKit)
Scripted turns in 14 languages
False-cutoff rate at 300 and 600 ms
LiveKit
Scripted
The Artificial Analysis index is one WER over three sets, half of it AA-AgentTalk, with no entity, noise, medical, multilingual, code-switching or diarization component; its latency runs from a voice-activity detector’s end of speech to the final transcript, a different quantity from the endpointing latency in §2. Coval scores against model-generated references. Neither substitutes for the other or for this report.
Caveats
We ran it. AssemblyAI produced the results in §1–§8. Competitors ran on default settings and may do better tuned. §9 has independent results.
Snapshots differ. Runs span 2026-08-18 to 2026-09-08 (Table 20). A vendor that updated its model within that window may appear on the older version.
Coverage is uneven. Not every system ran every set; means are withheld rather than computed on partial data.
Reference quality. TED-LIUM references omit sponsor read-outs. Some independent leaderboards (Coval) use model-generated references.
Latency is measured one way. Endpointing latency here runs from end of speech to the endpoint event. Independent leaderboards measure time to final transcript or to first token, which are not comparable.
AssemblyAI Research · Realtime benchmarks technical report · data as of 2026-09-09
Benchmarks
Run your own benchmark
Testing on your own audio? Want to do it right? Read our benchmarking guide in the docs.
Benchmark audio was evaluated with production model settings. No custom model tuning or prompt engineering was applied.
Transcription outputs were normalized using the Whisper text normalizer before WER computation. This removes formatting differences (casing, punctuation, number formats) so that scores reflect transcription quality, not output styling.
We evaluate across benchmark datasets spanning general speech, entity recognition, diarization, multilingual audio, and code switching.
Word error rate
Selected speech sets covering synthetic medical, accented English, general speech, and webinar audio.