New The Dictation API is here. Learn more: Dictation API
Benchmarks

AssemblyAI Universal-3.5 Pro benchmarks

Benchmarks for Universal-3.5 Pro across pre-recorded audio and Universal-3.5 Pro Realtime across realtime audio.

Technical report Realtime speech-to-text data as of 2026-09-09 v2026.09

Realtime speech recognition, measured against the field

Universal-3.5 Pro Realtime and 24 competitor models on 6 evaluation suites: same audio, same scoring pipeline.

Universal-3.5 Pro Realtime has the lowest entity error rate on the voice-agent suite: 14.78% against 15.95% for the next system, a 7% lead. It also leads on judge-scored entity accuracy in general English, speaker diarization, code-switched speech and end-of-turn recall.

Every score is in the tables. §9 reproduces independent evaluations; §10 describes the method.

At a glance

Voice-agent entity error
14.78%
1st of 17
next best ElevenLabs Scribe v2 15.95%
Judge-scored entity error
19.17%
1st of 17
next best ElevenLabs Scribe v2 21.19%
Diarization DER
17.62%
1st of 9
next best Azure 20.43%
Code-switching (5 pairs)
7.82%
1st of 16
next best AWS Transcribe 8.30%
Multilingual (FLEURS, 9 lang.)
7.20%
1st of 13
next best AWS Transcribe 7.37%

01Overview

Each cell is the vendor’s best model on that metric, coloured by rank within the column (solid blue is first). Values are in the tables of §2–§8.

Rank across 11 headline metrics, vendors ordered by mean rank.
Vendor (best model per cell) VA entity err. NEER % Entity (judge) % Medical MER % End-of-turn F1 % Endpoint P90 ms Diarization DER % Non-speech resp. % Code-switching WER % Multi-lingual FLEURS % EN short-form WER % EN long-form WER % meanrank
AssemblyAI 14.819.25.181.9162217.60.37.87.23.99.5 3.1
Speechmatics 27.824.84.728.1461822.20.238.112.44.77.7 5.9
ElevenLabs 15.921.285.519550.212.47.719.9 6.0
AWS 22.830.777.896527.51.38.37.48.0 6.0
Deepgram 33.830.713.276.674834.00.216.56.59.1 6.6
Mistral 23.123.412.326.739910.017.79.76.622.7 7.2
Azure 50.214.420.40.213.89.85.58.7 7.3
Google 21.68.684.11482 7.5
Smallest 46.840.721.474.0124741.41.413.112.05.49.0 7.6
xAI 26.936.115.770.795225.71.25.89.0 7.7
Reson8 25.315.00.76.9 8.8
OpenAI 36.725.956.810830.132.810.4 9.0
Soniox 20.424.01162531.510.8 10.2
Cartesia 35.929.860.95253.150.716.811.1 10.3
Gladia 28.04.916.314.111.7 12.2

Best Worst Cells are coloured by rank within each column (solid blue is best, white is worst).

Figure 1 Rank across 11 headline metrics, vendors ordered by mean rank. source = the tables below; each column’s provenance is given with its table

Position in the field

  • Voice-agent entity error 14.78% · 1st of 17 systems · runner-up ElevenLabs Scribe v2 15.95% −7.3%
  • Judge-scored entity error 19.17% · 1st of 17 systems · runner-up ElevenLabs Scribe v2 21.19% −9.5%
  • Medical entity error 5.10% · 2nd of 11 systems · leader Speechmatics Enhanced 4.66% +9.4%
  • End-of-turn F1 81.94% · 3rd of 17 systems · leader ElevenLabs Scribe v2 85.55% −3.6 pt
  • Endpointing P90 latency 1622 ms · 11th of 17 systems · leader Cartesia Ink-Whisper 525 ms +1097 ms
  • Diarization DER 17.62% · 1st of 9 systems · runner-up Azure 20.43% −13.8%
  • Non-speech response rate 0.27% · 8th of 17 systems · leader Mistral Voxtral Mini 0.03% +0.2 pt
  • Code-switching (5 pairs) 7.82% · 1st of 16 systems · runner-up AWS Transcribe 8.30% −5.8%
  • Multilingual (FLEURS, 9 lang.) 7.20% · 1st of 13 systems · runner-up AWS Transcribe 7.37% −2.3%
  • English short-form WER 3.87% · 1st of 9 systems · runner-up Speechmatics Enhanced 4.69% −17.5%
  • Long-form English WER 9.53% · 10th of 14 systems · leader Reson8 6.86% +38.9%

We lead We trail The margin is blue where we lead, grey where we trail.

One row per headline metric from the grid above: our value, our rank among systems with a complete run, and the margin to the runner-up where we lead or to the leader where we do not. Margins are relative for error rates and absolute for F1, latency, and rates below one percent.

02Voice agents

Conversational audio as a voice agent hears it, across 17 systems, three noise environments, 12 speaker groups and three utterance lengths. The headline metric, entity error (NEER), is how often the names, addresses and numbers an agent acts on come out wrong.

Table 1 Entity error rate (%) by entity class. Colour is rank within the column. We lead in 7 of 11 columns. normalized entity error rate (NEER) · runs 2026-08-27 → 2026-09-08 · one row per model, most recent run · corpus value = mean of the three natural-noise environments
Entity error rate (%) by entity class. Colour is rank within the column. We lead in 7 of 11 columns.
System Overall Names Places Addresses Phone Email Dates Numbers Codes/IDs Medical Technical
AssemblyAI Universal-3.5 Pro 14.811.16.127.92.815.721.115.410.419.126.6
ElevenLabs Scribe v2 15.97.38.829.63.317.627.933.612.812.319.9
Soniox Realtime 20.416.48.129.33.020.832.126.88.620.832.4
Google Gemini 3.5 Live 21.616.09.832.87.823.333.422.317.622.333.0
AWS Transcribe 22.821.87.430.23.721.030.624.512.821.540.0
Mistral Voxtral Mini 23.118.210.140.38.325.931.421.416.132.637.6
xAI Grok Streaming 26.919.012.885.013.729.824.638.626.437.042.5
Speechmatics Enhanced 27.824.47.037.585.622.733.921.122.918.538.1
Deepgram Nova-3 33.825.511.331.091.733.447.823.132.630.051.1
Cartesia Ink-Whisper 35.925.116.361.659.247.336.028.455.545.145.5
Speechmatics Standard 36.331.19.549.487.856.135.423.646.729.842.2
OpenAI GPT-4o Transcribe 36.728.218.356.244.149.636.134.656.338.045.7
OpenAI GPT-4o mini Transcribe 37.329.619.958.241.551.738.331.651.243.146.0
OpenAI GPT Realtime 39.432.623.353.339.145.644.036.452.837.645.6
Deepgram Flux EN 46.330.614.094.511.639.744.563.099.933.660.5
Smallest Pulse 46.830.417.094.613.238.748.355.496.544.653.5
Deepgram Flux Multi 47.633.714.294.610.136.944.463.399.938.859.5

Best Worst Cells are coloured by rank within each column (solid blue is best, white is worst).

Background noise and utterance length

0 10 20 30 40 AssemblyAI Universal-3.5 Pro · No background noise: 13.4% AssemblyAI Universal-3.5 Pro · Natural low noise: 14.4% AssemblyAI Universal-3.5 Pro · Natural heavy noise: 15.1% AssemblyAI Universal-3.5 Pro ElevenLabs Scribe v2 · No background noise: 14.4% ElevenLabs Scribe v2 · Natural low noise: 15.4% ElevenLabs Scribe v2 · Natural heavy noise: 16.2% ElevenLabs Scribe v2 Soniox Realtime · No background noise: 19.3% Soniox Realtime · Natural low noise: 20.0% Soniox Realtime · Natural heavy noise: 20.0% Soniox Realtime Google Gemini 3.5 Live · No background noise: 19.6% Google Gemini 3.5 Live · Natural low noise: 20.7% Google Gemini 3.5 Live · Natural heavy noise: 22.0% Google Gemini 3.5 Live Mistral Voxtral Mini · No background noise: 21.0% Mistral Voxtral Mini · Natural low noise: 22.2% Mistral Voxtral Mini · Natural heavy noise: 23.4% Mistral Voxtral Mini AWS Transcribe · No background noise: 21.4% AWS Transcribe · Natural low noise: 22.1% AWS Transcribe · Natural heavy noise: 23.1% AWS Transcribe xAI Grok Streaming · No background noise: 24.6% xAI Grok Streaming · Natural low noise: 25.8% xAI Grok Streaming · Natural heavy noise: 27.1% xAI Grok Streaming Speechmatics Enhanced · No background noise: 26.5% Speechmatics Enhanced · Natural low noise: 27.6% Speechmatics Enhanced · Natural heavy noise: 27.7% Speechmatics Enhanced Deepgram Nova-3 · No background noise: 31.9% Deepgram Nova-3 · Natural low noise: 33.0% Deepgram Nova-3 · Natural heavy noise: 34.1% Deepgram Nova-3 Cartesia Ink-Whisper · No background noise: 32.7% Cartesia Ink-Whisper · Natural low noise: 34.8% Cartesia Ink-Whisper · Natural heavy noise: 36.7% Cartesia Ink-Whisper OpenAI GPT-4o Transcribe · No background noise: 33.5% OpenAI GPT-4o Transcribe · Natural low noise: 35.8% OpenAI GPT-4o Transcribe · Natural heavy noise: 36.5% OpenAI GPT-4o Transcribe Speechmatics Standard · No background noise: 34.0% Speechmatics Standard · Natural low noise: 35.8% Speechmatics Standard · Natural heavy noise: 36.2% Speechmatics Standard OpenAI GPT-4o mini Transcribe · No background noise: 34.1% OpenAI GPT-4o mini Transcribe · Natural low noise: 36.8% OpenAI GPT-4o mini Transcribe · Natural heavy noise: 37.1% OpenAI GPT-4o mini Transcribe OpenAI GPT Realtime · No background noise: 35.3% OpenAI GPT Realtime · Natural low noise: 37.7% OpenAI GPT Realtime · Natural heavy noise: 39.3% OpenAI GPT Realtime Deepgram Flux EN · No background noise: 44.3% Deepgram Flux EN · Natural low noise: 45.4% Deepgram Flux EN · Natural heavy noise: 46.3% Deepgram Flux EN Smallest Pulse · No background noise: 45.3% Smallest Pulse · Natural low noise: 46.5% Smallest Pulse · Natural heavy noise: 44.9% Smallest Pulse Deepgram Flux Multi · No background noise: 45.4% Deepgram Flux Multi · Natural low noise: 46.4% Deepgram Flux Multi · Natural heavy noise: 47.9% Deepgram Flux Multi
No background noise Natural low noise Natural heavy noise
Figure 2 Entity error rate (%) by background-noise condition, sorted by the no-noise value. We lead in all 3 conditions. normalized entity error rate (NEER) · runs 2026-08-27 → 2026-09-08 · one row per model, most recent run

Listen to the test data

One caller turn per noise environment, quietest first, each with an entity an agent acts on: a name, an account number, an e-mail login, a medication. Words that differ from the human reference are marked.

1 of 8
Voice agentsno noise4 s · AssemblyAI voice-agent corpus — scripted caller scenarios voiced by real speakers

Four seconds, no noise. A short request where the other systems substitute unrelated words.

Reference

My cable modem keeps rebooting every 30-40 minutes, can you help?

AssemblyAI

My cable modem keeps rebooting every 30 to 40 minutes. Can you help?

ElevenLabs Scribe v2 Realtime

My key won't let them keep changing every 30 to 40 minutes. Can you help?

Soniox Realtime

Mark Kilmer and keeps removing every 30 to 40 minutes. Can you help?

0 10 20 30 40 50 AssemblyAI Universal-3.5 Pro · Short utterances: 13.8% AssemblyAI Universal-3.5 Pro · Medium: 13.7% AssemblyAI Universal-3.5 Pro · Long: 16.4% AssemblyAI Universal-3.5 Pro ElevenLabs Scribe v2 · Short utterances: 15.8% ElevenLabs Scribe v2 · Medium: 15.1% ElevenLabs Scribe v2 · Long: 17.2% ElevenLabs Scribe v2 Google Gemini 3.5 Live · Short utterances: 19.9% Google Gemini 3.5 Live · Medium: 20.5% Google Gemini 3.5 Live · Long: 25.2% Google Gemini 3.5 Live Soniox Realtime · Short utterances: 21.7% Soniox Realtime · Medium: 19.0% Soniox Realtime · Long: 19.7% Soniox Realtime AWS Transcribe · Short utterances: 23.3% AWS Transcribe · Medium: 21.8% AWS Transcribe · Long: 23.0% AWS Transcribe Mistral Voxtral Mini · Short utterances: 23.3% Mistral Voxtral Mini · Medium: 21.3% Mistral Voxtral Mini · Long: 24.1% Mistral Voxtral Mini xAI Grok Streaming · Short utterances: 25.9% xAI Grok Streaming · Medium: 25.7% xAI Grok Streaming · Long: 29.9% xAI Grok Streaming Speechmatics Enhanced · Short utterances: 25.9% Speechmatics Enhanced · Medium: 27.4% Speechmatics Enhanced · Long: 29.9% Speechmatics Enhanced Deepgram Nova-3 · Short utterances: 32.0% Deepgram Nova-3 · Medium: 33.8% Deepgram Nova-3 · Long: 36.1% Deepgram Nova-3 OpenAI GPT Realtime · Short utterances: 32.8% OpenAI GPT Realtime · Medium: 38.0% OpenAI GPT Realtime · Long: 49.9% OpenAI GPT Realtime OpenAI GPT-4o Transcribe · Short utterances: 33.5% OpenAI GPT-4o Transcribe · Medium: 36.0% OpenAI GPT-4o Transcribe · Long: 41.1% OpenAI GPT-4o Transcribe Speechmatics Standard · Short utterances: 33.8% Speechmatics Standard · Medium: 36.0% Speechmatics Standard · Long: 38.9% Speechmatics Standard OpenAI GPT-4o mini Transcribe · Short utterances: 34.1% OpenAI GPT-4o mini Transcribe · Medium: 36.4% OpenAI GPT-4o mini Transcribe · Long: 41.6% OpenAI GPT-4o mini Transcribe Cartesia Ink-Whisper · Short utterances: 34.3% Cartesia Ink-Whisper · Medium: 35.1% Cartesia Ink-Whisper · Long: 38.1% Cartesia Ink-Whisper Smallest Pulse · Short utterances: 43.1% Smallest Pulse · Medium: 46.4% Smallest Pulse · Long: 51.0% Smallest Pulse Deepgram Flux EN · Short utterances: 44.4% Deepgram Flux EN · Medium: 45.4% Deepgram Flux EN · Long: 49.3% Deepgram Flux EN Deepgram Flux Multi · Short utterances: 46.0% Deepgram Flux Multi · Medium: 47.0% Deepgram Flux Multi · Long: 50.4% Deepgram Flux Multi
Short utterances Medium Long
Figure 3 Entity error rate (%) by utterance length. We lead in all 3 length bands. normalized entity error rate (NEER) · runs 2026-08-27 → 2026-09-08 · one row per model, most recent run

Speaker groups

Table 2 Entity error rate (%) by speaker group. Colour is rank within the column. We lead in 10 of 12 groups. normalized entity error rate (NEER) · runs 2026-08-27 → 2026-09-08 · one row per model, most recent run
Entity error rate (%) by speaker group. Colour is rank within the column. We lead in 10 of 12 groups.
System US Northeast US South US Gen. Am. US AAVE US Asian-Am. US other UK Canada Native other Spanish-acc. Indian-acc. Non-nat. other meanrank
AssemblyAI Universal-3.5 Pro 10.912.712.414.013.815.815.111.716.517.321.720.2 1.2
ElevenLabs Scribe v2 13.415.414.215.414.515.617.013.617.018.820.720.9 1.8
Soniox Realtime 17.618.418.219.620.019.821.817.521.924.925.726.3 3.2
Google Gemini 3.5 Live 18.921.419.521.718.522.623.918.422.826.023.327.5 4.2
Mistral Voxtral Mini 18.921.919.821.821.224.122.817.524.930.330.032.4 5.3
AWS Transcribe 19.621.721.022.621.522.722.919.824.326.326.928.8 5.4
xAI Grok Streaming 21.826.924.626.623.928.228.321.930.529.833.533.5 7.3
Speechmatics Enhanced 25.726.925.526.326.728.827.925.429.131.731.433.9 7.7
Deepgram Nova-3 30.733.531.932.932.233.134.530.535.437.938.639.9 9.3
Cartesia Ink-Whisper 30.335.832.334.534.038.037.327.638.044.542.346.1 10.5
Speechmatics Standard 32.334.933.136.136.237.936.833.638.242.040.943.8 11.6
OpenAI GPT-4o Transcribe 31.437.332.935.433.238.138.728.438.546.142.447.4 11.8
OpenAI GPT-4o mini Transcribe 31.537.433.035.134.139.139.229.439.446.544.648.9 12.7
OpenAI GPT Realtime 34.936.634.934.535.537.840.934.743.048.848.151.6 13.2
Deepgram Flux EN 43.945.743.646.543.544.548.644.148.750.550.452.1 15.3
Smallest Pulse 44.644.543.448.144.549.747.740.448.753.053.554.7 16.1
Deepgram Flux Multi 45.246.745.148.244.445.849.745.150.252.951.954.1 16.6

Best Worst Cells are coloured by rank within each column (solid blue is best, white is worst).

Turn-taking

Table 3 End-of-turn detection and endpointing latency. Higher is better for F1, precision and recall; lower for latency. We lead on recall. end-of-turn detection (F1/precision/recall, %) and endpointing latency (ms) · runs 2026-08-27 → 2026-09-08 · one row per model, most recent run
End-of-turn detection and endpointing latency. Higher is better for F1, precision and recall; lower for latency. We lead on recall.
# System model / snapshot F1 % Precision % Recall % P50 ms P90 ms P99 ms Mean ms
1 ElevenLabs Scribe v2 Realtime 85.583.387.91775195523911640
2 Google Gemini 3.5 Transcribe Live 84.181.587.01306148221431254
3 AssemblyAI Universal-3.5 Pro Realtime 81.973.892.154616222909742
4 AWS Transcribe Streaming 77.869.488.47659651653728
5 Deepgram Flux Multi 76.668.686.83287481479391
6 Deepgram Flux EN 75.468.184.53488271672402
7 Smallest Pulse 74.063.588.992512472351954
8 Deepgram Nova-3 71.559.390.33978495186572
9 xAI Grok Streaming 70.760.685.17079521554688
10 Cartesia Ink-Whisper 60.948.482.5330525747291
11 OpenAI GPT-4o Transcribe 56.844.479.087710871644847
12 OpenAI GPT-4o mini Transcribe 55.943.977.288010831645841
13 Speechmatics Enhanced Realtime 28.116.691.94410461848043885
14 Speechmatics Standard Realtime 26.815.988.54543475249063947
15 Mistral Voxtral Mini Realtime 26.791.616.28163991130852061
16 OpenAI GPT Realtime 24.684.015.01740310918951010124
17 Soniox Realtime 24.017.239.9801611625142096501

Best Others The best value in each column is filled blue.

500 1,000 2,000 5,000 10,000 20,000 20% 40% 60% 80% Endpointing latency P90 (ms, lower is better) End-of-turn F1 (%, higher is better) Soniox Realtime — 11,625 ms · 24.02% Soniox Realtime 11,625 ms · 24.02% OpenAI GPT Realtime — 31,091 ms · 24.60% OpenAI GPT Realtime 31,091 ms · 24.60% Mistral Voxtral Mini — 3,991 ms · 26.66% Mistral Voxtral Mini 3,991 ms · 26.66% Speechmatics Standard — 4,752 ms · 26.84% Speechmatics Standard 4,752 ms · 26.84% Speechmatics Enhanced — 4,618 ms · 28.05% Speechmatics Enhanced 4,618 ms · 28.05% OpenAI GPT-4o mini Transcribe — 1,083 ms · 55.92% OpenAI GPT-4o mini Transcribe 1,083 ms · 55.92% OpenAI GPT-4o Transcribe — 1,087 ms · 56.78% OpenAI GPT-4o Transcribe 1,087 ms · 56.78% Cartesia Ink-Whisper — 525 ms · 60.91% Cartesia Ink-Whisper 525 ms · 60.91% xAI Grok Streaming — 952 ms · 70.72% xAI Grok Streaming 952 ms · 70.72% Deepgram Nova-3 — 849 ms · 71.51% Deepgram Nova-3 849 ms · 71.51% Smallest Pulse — 1,247 ms · 74.02% Smallest Pulse 1,247 ms · 74.02% Deepgram Flux EN — 827 ms · 75.42% Deepgram Flux EN 827 ms · 75.42% Deepgram Flux Multi — 748 ms · 76.60% 1 AWS Transcribe — 965 ms · 77.75% AWS Transcribe 965 ms · 77.75% AssemblyAI Universal-3.5 Pro — 1,622 ms · 81.94% AssemblyAI Universal-3.5 Pro 1,622 ms · 81.94% Google Gemini 3.5 Live — 1,482 ms · 84.14% 2 ElevenLabs Scribe v2 — 1,955 ms · 85.55% ElevenLabs Scribe v2 1,955 ms · 85.55%
  1. 1 Deepgram Flux Multi · 748 ms · 76.60%
  2. 2 Google Gemini 3.5 Live · 1,482 ms · 84.14%
Figure 4 End-of-turn F1 against P90 endpointing latency (log scale). Upper left is better. end-of-turn detection (F1/precision/recall, %) and endpointing latency (ms) · runs 2026-08-27 → 2026-09-08 · one row per model, most recent run

Ablations on our own model

wordboost OFF wordboost ON
Figure 5 Word boost (custom vocabulary) on versus off, same model and corpus. Change in entity error with word boost on: names −58.5%, technical terms −68.4%. normalized entity error rate (NEER) + WER, wordboost OFF vs ON · runs 2026-08-27 → 2026-09-03 · one row per model, most recent run
agent context OFF agent context ON
Figure 6 Agent context on versus off: the agent’s previous turn is passed to the recognizer. normalized entity error rate (NEER) + WER, same model with and without agent context · runs 2026-08-27 → 2026-08-28 · one row per model, most recent run

Word error rate

Table 4 WER by noise condition. normalized word error rate · runs 2026-08-27 → 2026-09-08 · one row per model, most recent run
WER by noise condition.
# System model / snapshot Mean 3 conditions No noise Low noise Heavy noise
1 AssemblyAI Universal-3.5 Pro Realtime 5.81%4.63%5.60%7.19%
2 Soniox Realtime 6.84%5.96%6.86%7.70%
3 ElevenLabs Scribe v2 Realtime 6.97%5.59%6.79%8.52%
4 Speechmatics Enhanced Realtime 8.14%6.96%7.96%9.49%
5 Google Gemini 3.5 Transcribe Live 8.48%7.44%8.49%9.52%
6 AWS Transcribe Streaming 8.58%7.27%8.27%10.19%
7 Deepgram Nova-3 8.64%7.53%8.45%9.93%
8 Mistral Voxtral Mini Realtime 10.43%9.25%10.45%11.59%
9 Smallest Pulse 10.74%9.85%10.94%11.43%
10 Speechmatics Standard Realtime 12.05%10.47%11.77%13.91%
11 Deepgram Flux EN 13.50%12.29%13.36%14.85%
12 Deepgram Flux Multi 14.26%12.85%14.06%15.87%
13 Cartesia Ink-Whisper 15.17%13.71%15.00%16.81%
14 OpenAI GPT-4o Transcribe 18.37%17.05%18.44%19.63%
15 OpenAI GPT-4o mini Transcribe 19.55%18.12%19.77%20.76%
16 xAI Grok Streaming 20.28%18.45%19.89%22.50%
17 OpenAI GPT Realtime 28.37%25.45%28.14%31.53%

Best Others The best value in each column is filled blue.

Data

Caller corpus 12,460 clips
All entity classes pooled into one number: how often an agent-critical value (a name, address, number or identifier) comes back wrong.
Noise environments
Three levels of real noise, recorded rather than mixed in. None: room and handset noise only. Low: incidental noise from a home, office or parked car. Heavy: noise that competes with the voice.
Utterance length
Short turns (confirmations, single values, one-word answers), medium and long. Short turns give the model almost no context.
Speaker groups
Fourteen groups: US Northeast, South, General American, AAVE, Asian-American and other; UK; Canada; native speakers overall; Spanish- and Indian-accented and other non-native speakers.
Entity classes
Names (including letter-by-letter spellings), places, street addresses, phone numbers, email addresses, dates and times, numbers and currency, codes and identifiers, medical and technical terms.

public, open source public, licensed private, ours

03Healthcare

Medical entity error is measured on primary-care consultations (PriMock57), licensed clinical dictation and respiratory-clinic recordings.

Overall

Table 5 Medical entity error rate. We lead on primary care. normalized medical entity error rate (MER) · runs 2026-08-18 → 2026-09-02 · one row per model, most recent run
Medical entity error rate. We lead on primary care.
# System model / snapshot Mean 3 sets Primary care PriMock57 Dictation Respiratory
1 Speechmatics Enhanced Realtime 4.66%8.13%3.65%2.21%
2 AssemblyAI Universal-3.5 Pro Realtime 5.10%7.18%5.23%2.90%
3 Google Gemini 3.5 Transcribe Live 8.57%15.45%7.69%2.56%
4 Mistral Voxtral Mini Realtime 12.26%12.67%16.02%8.09%
5 Deepgram Nova-3 13.22%19.01%17.37%3.29%
6 Speechmatics Standard Realtime 14.24%18.68%18.72%5.32%
7 Deepgram Flux EN 14.28%20.49%18.32%4.02%
8 Azure Realtime STT 14.41%17.21%19.11%6.92%
9 Reson8 Reson8 15.02%18.49%21.81%4.76%
10 xAI Grok Streaming 15.69%17.95%19.35%9.78%
11 Smallest Pulse 21.40%25.42%30.77%8.01%
AWS Transcribe Streaming incomplete11.09%16.49%
Cartesia Ink-Whisper incomplete24.39%31.09%
ElevenLabs Scribe v2 Realtime incomplete10.89%8.01%
Gladia Solaria incomplete18.83%25.30%
OpenAI GPT-4o mini Transcribe incomplete25.45%19.27%
OpenAI GPT-4o Transcribe incomplete23.34%14.91%
OpenAI GPT Realtime incomplete13.73%5.39%

Best Others The best value in each column is filled blue.

Medication and disease

Table 6 Medication and disease terms scored separately. We lead on the disease mean. medication / disease entity error rate · runs 2026-08-18 → 2026-09-02 · one row per model, most recent run
Medication and disease terms scored separately. We lead on the disease mean.
# System model / snapshot Medication mean Primary care Dictation Respiratory Disease mean Primary care Dictation Respiratory
1 Speechmatics Enhanced Realtime 6.62%8.92%5.04%5.90%3.38%7.35%0.93%1.86%
2 AssemblyAI Universal-3.5 Pro Realtime 8.70%9.55%7.31%9.25%2.52%4.83%1.17%1.57%
3 Google Gemini 3.5 Transcribe Live 12.23%19.19%10.80%6.70%5.16%11.76%1.64%2.09%
4 Deepgram Nova-3 17.76%25.05%20.53%7.69%8.83%13.03%11.21%2.24%
5 Mistral Voxtral Mini Realtime 18.22%18.47%22.09%14.10%5.98%6.93%4.21%6.79%
6 Azure Realtime STT 18.60%23.57%21.97%10.26%9.57%10.92%13.55%4.25%
7 xAI Grok Streaming 19.50%24.63%26.05%7.81%7.74%11.34%6.31%5.56%
8 Deepgram Flux EN 19.77%29.30%20.77%9.23%9.18%11.76%13.55%2.24%
9 Speechmatics Standard Realtime 19.81%26.98%21.13%11.31%9.56%10.53%14.02%4.13%
10 Reson8 Reson8 22.55%27.74%27.85%12.05%7.15%8.80%10.05%2.61%
11 Smallest Pulse 26.83%31.55%41.42%7.53%12.04%18.30%10.05%7.76%
AWS Transcribe Streaming incomplete15.07%21.97%incomplete7.14%5.84%
Cartesia Ink-Whisper incomplete33.55%37.01%incomplete15.34%19.57%
ElevenLabs Scribe v2 Realtime incomplete13.57%11.64%incomplete8.17%0.93%
Gladia Solaria incomplete23.81%32.41%incomplete12.86%11.45%
OpenAI GPT-4o mini Transcribe incomplete32.91%21.49%incomplete18.07%14.95%
OpenAI GPT-4o Transcribe incomplete27.18%16.93%incomplete19.54%10.98%
OpenAI GPT Realtime incomplete17.83%6.72%incomplete9.66%2.80%

Best Others The best value in each column is filled blue.

Data

Primary-care consultations (PriMock57) 114 consultations
Public corpus of mock primary-care consultations: clinicians and actor-patients work through realistic appointments, recorded as telemedicine calls.
Medical dictation 93 recordings
Licensed clinician dictation, notes and findings spoken for the record, dense with medication and disease terms.
Respiratory clinic recordings 272 recordings
Real clinical conversation from respiratory clinic encounters, with the accents, interruptions and room noise of a working clinic.

public, open source public, licensed private, ours

04English

General English: entity accuracy and non-speech handling first, then word error rate on short-form, accented and long-form sets.

Table 7 Judge-scored entity error (%) on general English, by class. We lead on both overall scores. Per class, we lead on person names, product names, street addresses, organisations and URLs. A 0.00 cell means no data. LLM-judge-scored entity error rate (normalized and formatted) · runs 2026-08-18 → 2026-09-02 · one row per model, most recent run
Judge-scored entity error (%) on general English, by class. We lead on both overall scores. Per class, we lead on person names, product names, street addresses, organisations and URLs. A 0.00 cell means no data.
System All classes All, formatted Persons Products Addresses Orgs Places Phone Email URLs Dates Credit cards Alphanum.
AssemblyAI Universal-3.5 Pro 19.227.819.920.924.018.421.314.320.814.221.813.78.0
ElevenLabs Scribe v2 21.229.224.825.532.623.219.1n/d19.914.922.0n/d9.2
Mistral Voxtral Mini 23.431.827.828.028.919.617.53.895.531.019.9n/d7.8
Speechmatics Enhanced 24.837.725.326.836.622.216.886.569.723.521.1n/d9.2
Reson8 25.338.437.027.231.423.618.58.016.326.022.94.29.3
OpenAI GPT-4o Transcribe 25.938.126.228.641.323.426.418.139.419.724.821.413.9
OpenAI GPT-4o mini Transcribe 26.638.628.729.441.324.227.38.945.720.423.28.314.4
OpenAI GPT Realtime 27.638.724.930.337.620.218.02.198.252.522.73.044.3
Gladia Solaria 28.041.832.530.438.424.322.44.295.932.324.17.711.3
Cartesia Ink-Whisper 29.843.337.730.539.427.626.312.753.425.424.813.712.8
Deepgram Nova-3 30.743.536.434.037.627.022.781.922.622.240.27.118.0
AWS Transcribe 30.744.146.233.626.828.223.90.479.628.722.98.37.7
Speechmatics Standard 32.344.634.029.941.326.820.685.792.341.022.1n/d36.9
xAI Grok Streaming 36.151.834.234.255.626.126.92.598.280.857.24.812.4
Smallest Pulse 40.758.453.237.271.832.430.23.823.125.324.85.461.4
Azure 50.297.546.042.089.636.830.52.198.272.966.30.6100.0
Deepgram Flux EN 50.562.549.439.388.231.726.33.498.299.163.1n/d100.0

Best Worst Cells are coloured by rank within each column (solid blue is best, white is worst).

0 2 4 Mistral Voxtral Mini: 0.03% Mistral Voxtral Mini 0.03% OpenAI GPT Realtime: 0.07% OpenAI GPT Realtime 0.07% Deepgram Nova-3: 0.17% Deepgram Nova-3 0.17% ElevenLabs Scribe v2: 0.20% ElevenLabs Scribe v2 0.20% Speechmatics Standard: 0.20% Speechmatics Standard 0.20% Azure: 0.24% Azure 0.24% Speechmatics Enhanced: 0.27% Speechmatics Enhanced 0.27% AssemblyAI Universal-3.5 Pro: 0.27% AssemblyAI Universal-3.5 Pro 0.27% OpenAI GPT-4o mini Transcribe: 0.37% OpenAI GPT-4o mini Transcribe 0.37% Reson8: 0.68% Reson8 0.68% xAI Grok Streaming: 1.22% xAI Grok Streaming 1.22% AWS Transcribe: 1.25% AWS Transcribe 1.25% Smallest Pulse: 1.42% Smallest Pulse 1.42% OpenAI GPT-4o Transcribe: 2.07% OpenAI GPT-4o Transcribe 2.07% Cartesia Ink-Whisper: 3.15% Cartesia Ink-Whisper 3.15% Gladia Solaria: 4.88% Gladia Solaria 4.88% Deepgram Flux EN: 38.75% Deepgram Flux EN › 38.75%
Figure 7 Non-speech response rate: how often a system produces text for a clip with no speech. Lower is better. The axis is clipped at 5%; the notched bar exceeds it. non-speech response rate (share of pure non-speech clips that produce any text) · runs 2026-08-18 → 2026-09-02 · one row per model, most recent run

Word error rate — short-form

Table 8 Short-form English WER. We lead on the mean and Common Voice. Systems without a Common Voice run have no mean. normalized word error rate · runs 2026-08-18 → 2026-09-02 · one row per model, most recent run
Short-form English WER. We lead on the mean and Common Voice. Systems without a Common Voice run have no mean.
# System model / snapshot Mean 3 sets Common Voice LibriSpeech clean LibriSpeech other
1 AssemblyAI Universal-3.5 Pro Realtime 3.87%6.44%1.88%3.28%
2 Speechmatics Enhanced Realtime 4.69%6.78%2.31%4.99%
3 Smallest Pulse 5.41%9.34%2.19%4.69%
4 Azure Realtime STT 5.54%8.68%2.44%5.49%
5 xAI Grok Streaming 5.81%11.04%2.04%4.36%
6 Speechmatics Standard Realtime 6.41%9.43%3.18%6.61%
7 Deepgram Flux EN 6.53%7.68%3.56%8.35%
8 Mistral Voxtral Mini Realtime 6.61%12.23%2.10%5.49%
9 Deepgram Nova-3 7.46%12.38%3.28%6.72%
AWS Transcribe Streaming incomplete1.90%3.77%
Cartesia Ink-Whisper incomplete3.19%6.48%
ElevenLabs Scribe v2 Realtime incomplete1.96%4.43%
Gladia Solaria incomplete2.73%5.42%
OpenAI GPT-4o mini Transcribe incomplete2.53%7.97%
OpenAI GPT-4o Transcribe incomplete2.20%7.15%
OpenAI GPT Realtime incomplete2.59%7.44%
Reson8 Reson8 incomplete1.44%2.68%

Best Others The best value in each column is filled blue.

Listen to the test data

Two read sentences from the short-form sets above, human reference beside each transcript. Words that differ from it are marked.

1 of 2
Short-form English17 s · Common Voice (CC0)

A read sentence. The other system appends a sentence and a name that are not in the audio.

Reference

You can do a lot just using your voice, but there are still a few times you'll find yourself reaching for a mouse.

AssemblyAI

You can do a lot just using your voice. But there are still a few times you'll find yourself reaching for a mouse.

Speechmatics Enhanced

I know you can do a lot just using your voice, but there are still a few times you'll find yourself reaching for a mouse. And I'm Larry Hanson. We're here in this beautiful hall.

Word error rate — accented English

Table 9 Accented English WER. We lead on British-accented English and Indian-accented English. normalized word error rate · runs 2026-08-18 → 2026-09-02 · one row per model, most recent run
Accented English WER. We lead on British-accented English and Indian-accented English.
# System model / snapshot Mean 4 sets British French-acc. Spanish-acc. Indian-acc.
1 Azure Realtime STT 6.09%6.25%4.44%5.52%8.15%
2 AssemblyAI Universal-3.5 Pro Realtime 7.47%5.67%8.67%9.60%5.94%
3 Reson8 Reson8 9.25%5.76%11.89%12.60%6.73%
4 xAI Grok Streaming 9.85%7.33%12.21%12.88%6.99%
5 ElevenLabs Scribe v2 Realtime 10.10%6.92%12.12%13.29%8.07%
6 OpenAI GPT-4o Transcribe 10.30%7.27%11.83%13.89%8.22%
7 Speechmatics Enhanced Realtime 10.44%6.70%13.26%14.80%6.98%
8 AWS Transcribe Streaming 11.10%7.02%15.49%15.47%6.42%
9 Cartesia Ink-Whisper 11.23%8.16%12.92%14.57%9.27%
10 OpenAI GPT-4o mini Transcribe 11.36%7.88%13.31%15.35%8.88%
11 OpenAI GPT Realtime 11.78%8.31%14.83%15.78%8.20%
12 Smallest Pulse 11.85%7.70%15.15%16.53%8.03%
13 Deepgram Nova-3 11.96%9.14%13.92%15.19%9.57%
14 Deepgram Flux EN 12.55%9.61%14.83%16.12%9.64%
15 Speechmatics Standard Realtime 12.58%8.43%16.29%17.46%8.14%
16 Mistral Voxtral Mini Realtime 12.63%7.39%14.19%15.87%13.06%
Gladia Solaria incomplete8.67%12.60%13.85%

Best Others The best value in each column is filled blue.

Word error rate — long-form

Table 10 Long-form English WER. TED-LIUM references omit sponsor read-outs present in the audio. normalized word error rate · runs 2026-08-18 → 2026-09-02 · one row per model, most recent run
Long-form English WER. TED-LIUM references omit sponsor read-outs present in the audio.
# System model / snapshot Mean 3 sets TED-LIUM Rev16 CALLHOME per channel
1 Reson8 Reson8 6.86%4.63%6.87%9.09%
2 Speechmatics Enhanced Realtime 7.74%6.29%7.34%9.59%
3 AWS Transcribe Streaming 8.03%6.45%7.81%9.82%
4 Speechmatics Standard Realtime 8.55%6.83%8.06%10.77%
5 Azure Realtime STT 8.73%7.16%8.17%10.87%
6 Smallest Pulse 8.99%5.23%7.62%14.13%
7 xAI Grok Streaming 9.04%6.30%8.00%12.82%
8 Deepgram Flux EN 9.07%7.01%8.08%12.12%
9 Deepgram Nova-3 9.34%7.43%9.40%11.19%
10 AssemblyAI Universal-3.5 Pro Realtime 9.53%8.63%8.36%11.61%
11 Cartesia Ink-Whisper 11.06%7.13%10.42%15.62%
12 Gladia Solaria 11.68%7.75%9.04%18.25%
13 ElevenLabs Scribe v2 Realtime 19.95%5.96%8.38%45.51%
14 Mistral Voxtral Mini Realtime 22.73%6.96%8.33%52.91%
OpenAI GPT-4o mini Transcribe incomplete7.28%21.26%
OpenAI GPT-4o Transcribe incomplete9.88%28.51%
OpenAI GPT Realtime incomplete5.70%6.02%

Best Others The best value in each column is filled blue.

Data

Judge-scored entities
The short-form English audio below. An LLM judge grades each entity (person, product, organisation, place, address, phone, email, date, ID, URL) on whether the transcript names the same real-world thing, with and without formatting.
Non-speech 2,950 clips
Hold music, line noise and silence, with no words at all. The score is how often a model emits text anyway.
Common Voice 16,029 clips
Crowdsourced read sentences from Mozilla’s public Common Voice, recorded by volunteers on their own phones and laptops, so accents, microphones and rooms vary widely.
LibriSpeech test-clean 2,620 utterances
Read English audiobook speech from the public LibriSpeech corpus, recorded close-mic by LibriVox volunteers. The "clean" split is the easier half: clear speakers, little noise.
LibriSpeech test-other 2,939 utterances
The harder half of the same LibriSpeech set: speakers and recordings that automatic transcription finds more difficult, still read prose.
Accented English 9,315 files
Four licensed sets of prompted British-, French-, Spanish- and Indian-accented English: accuracy outside the General American accent most models are tuned on.
TED-LIUM 11 full-length talks
Full TED talks from the public TED-LIUM corpus: one prepared speaker on stage, recorded end to end, many minutes of continuous speech per file.
Rev16 9 full-length recordings
Podcast-style recordings from the public Rev16 set: spontaneous multi-speaker studio conversation, episode length.
CALLHOME 71 scored channels
Public CALLHOME telephone calls between friends and family: unscripted, fast, overlapping conversation over a real phone line. Each speaker's channel is scored on its own.

public, open source public, licensed private, ours

05Multilingual

Nine languages on FLEURS and Common Voice; every system is compared on the same nine.

FLEURS

Table 11 WER on FLEURS (CER for languages marked *). The nine-language mean is shown only for complete runs. normalized word error rate (character error rate for Japanese, Chinese, Korean) · runs 2026-08-25 → 2026-09-03 · one row per model, most recent run
WER on FLEURS (CER for languages marked *). The nine-language mean is shown only for complete runs.
System Mean 9 lang., excl. ko Spanish French German Italian Portuguese Dutch Hindi Japanese* Chinese* Korean*
AssemblyAI Universal-3.5 Pro 7.24.35.25.34.55.76.712.56.713.9n/s
AWS Transcribe 7.45.37.68.04.68.17.16.37.012.34.9
ElevenLabs Scribe v2 7.74.66.95.84.36.47.015.88.610.14.6
Mistral Voxtral Mini 9.74.89.67.85.56.99.516.89.616.55.8
Azure 9.87.711.49.18.310.512.98.68.611.59.4
OpenAI GPT Realtime 10.45.711.811.25.98.810.110.413.416.37.1
OpenAI GPT-4o Transcribe 11.24.310.613.15.58.513.012.416.117.46.7
Smallest Pulse 12.08.912.711.67.011.313.110.712.720.17.4
Speechmatics Enhanced 12.44.76.56.85.47.38.57.031.333.85.4
Speechmatics Standard 13.05.48.18.25.77.39.57.031.434.75.2
OpenAI GPT-4o mini Transcribe 13.35.513.214.96.910.916.816.815.419.37.6
Gladia Solaria 14.17.312.710.97.58.711.623.923.621.16.6
Cartesia Ink-Whisper 16.89.016.912.79.710.717.427.520.426.310.9
Soniox Realtime incomplete5.07.86.87.64.4
Deepgram Nova-3 incomplete8.914.013.310.013.314.622.318.8n/sn/s
Deepgram Flux incomplete8.113.910.49.212.716.420.514.4n/sn/s

Best Worst Cells are coloured by rank within each column (solid blue is best, white is worst).

Common Voice

Table 12 WER on Common Voice (CER for languages marked *). normalized word error rate (character error rate for Japanese, Chinese, Korean) · runs 2026-08-25 → 2026-09-03 · one row per model, most recent run
WER on Common Voice (CER for languages marked *).
System Mean 9 lang., excl. ko Spanish French German Italian Portuguese Dutch Hindi Japanese* Chinese* Korean*
AWS Transcribe 6.64.18.85.84.24.62.75.115.98.57.5
Speechmatics Enhanced 8.72.57.23.74.75.22.43.923.225.17.2
AssemblyAI Universal-3.5 Pro 8.74.48.45.04.58.13.710.619.913.7n/s
Speechmatics Standard 10.14.210.26.25.36.43.04.025.425.88.3
OpenAI GPT Realtime 10.96.912.69.37.89.46.88.420.216.410.5
ElevenLabs Scribe v2 11.05.810.27.16.19.75.521.620.712.111.0
Smallest Pulse 12.88.612.311.36.621.69.29.717.518.010.7
OpenAI GPT-4o Transcribe 13.28.915.510.39.213.27.810.022.021.412.1
Mistral Voxtral Mini 13.75.711.58.87.49.98.518.122.830.514.3
OpenAI GPT-4o mini Transcribe 15.49.816.811.710.614.311.921.522.319.312.2
Cartesia Ink-Whisper 17.610.417.712.411.111.59.727.928.329.213.8
Soniox Realtime incomplete6.412.27.67.28.64.619.214.18.7
Azure incomplete6.811.06.16.84.64.919.89.910.4
Gladia Solaria incomplete8.114.410.59.89.522.026.229.09.6
Deepgram Nova-3 incomplete8.413.812.810.812.79.325.229.7n/sn/s
Deepgram Flux incomplete9.114.211.013.912.511.223.429.9n/sn/s

Best Worst Cells are coloured by rank within each column (solid blue is best, white is worst).

Data

FLEURS
Public corpus of parallel sentences read by native speakers in quiet settings, one test set per language. Shown: Spanish, French, German, Italian, Portuguese, Dutch, Hindi, Japanese and Chinese.
Common Voice
Mozilla’s public Common Voice, per language: volunteers reading sentences on their own devices, so accents and microphones vary. Japanese and Chinese are scored by character error rate.

public, open source public, licensed private, ours

06Code-switching

Speech that switches between English and a second language within one utterance: five language pairs plus two Spanish–English conversational corpora (Bangor Miami).

Table 13 Code-switching WER. We lead on the five-pair mean. We lead in 4 of 5 pairs. normalized word error rate · runs 2026-08-18 → 2026-09-06 · one row per model, most recent run · pair value = mean of 4 sub-sets
Code-switching WER. We lead on the five-pair mean. We lead in 4 of 5 pairs.
# System model / snapshot Mean 5 pairs EN↔ES EN↔DE EN↔FR EN↔IT EN↔PT Bangor Miami Miami 10h
1 AssemblyAI Universal-3.5 Pro Realtime 7.82%6.80%7.84%8.54%7.21%8.71%31.65%27.31%
2 AWS Transcribe Streaming 8.30%7.17%8.12%9.33%7.05%9.81%31.65%27.96%
3 Soniox Realtime 10.77%9.65%10.82%12.11%10.09%11.16%26.84%21.97%
4 ElevenLabs Scribe v2 Realtime 12.37%11.06%12.29%13.16%11.53%13.79%27.77%27.09%
5 Smallest Pulse 13.13%11.40%13.14%13.61%10.79%16.73%43.62%47.17%
6 Azure Realtime STT 13.75%13.02%13.33%14.86%12.52%15.03%34.83%34.17%
7 Gladia Solaria 16.28%14.36%18.10%18.73%14.28%15.93%35.50%34.49%
8 Deepgram Nova-3 Multi 16.47%13.36%18.01%18.52%15.17%17.29%31.04%29.91%
9 Mistral Voxtral Mini Realtime 17.67%14.40%19.11%20.79%16.61%17.43%34.76%34.35%
10 OpenAI GPT-4o mini Transcribe 32.78%32.84%33.06%34.22%31.73%32.07%48.24%45.01%
11 OpenAI GPT-4o Transcribe 34.38%34.26%34.70%35.78%33.17%34.00%61.76%61.27%
12 OpenAI GPT Realtime 36.65%35.85%35.46%37.87%36.30%37.77%55.84%53.31%
13 Speechmatics Enhanced Realtime 38.11%33.18%33.78%35.98%45.39%42.23%43.51%36.04%
14 Speechmatics Standard Realtime 42.70%36.77%44.12%42.99%45.11%44.49%48.14%40.95%
15 Cartesia Ink-Whisper EN 50.69%47.71%51.50%53.10%53.17%47.95%52.94%43.72%
16 Deepgram Nova-3 EN 55.23%57.02%54.61%58.15%56.24%50.15%52.21%43.85%

Best Others The best value in each column is filled blue.

0 20 40 AssemblyAI Universal-3.5 Pro: 7.82% AssemblyAI Universal-3.5 Pro 7.82% AWS Transcribe: 8.30% AWS Transcribe 8.30% Soniox Realtime: 10.77% Soniox Realtime 10.77% ElevenLabs Scribe v2: 12.37% ElevenLabs Scribe v2 12.37% Smallest Pulse: 13.13% Smallest Pulse 13.13% Azure: 13.75% Azure 13.75% Gladia Solaria: 16.28% Gladia Solaria 16.28% Deepgram Nova-3 Multi: 16.47% Deepgram Nova-3 Multi 16.47% Mistral Voxtral Mini: 17.67% Mistral Voxtral Mini 17.67% OpenAI GPT-4o mini Transcribe: 32.78% OpenAI GPT-4o mini Transcribe 32.78% OpenAI GPT-4o Transcribe: 34.38% OpenAI GPT-4o Transcribe 34.38% OpenAI GPT Realtime: 36.65% OpenAI GPT Realtime 36.65% Speechmatics Enhanced: 38.11% Speechmatics Enhanced 38.11% Speechmatics Standard: 42.70% Speechmatics Standard 42.70% Cartesia Ink-Whisper EN: 50.69% Cartesia Ink-Whisper EN 50.69% Deepgram Nova-3 EN: 55.23% Deepgram Nova-3 EN 55.23%
Figure 8 Five-pair mean WER. normalized word error rate · runs 2026-08-18 → 2026-09-06 · one row per model, most recent run

Data

English ↔ Spanish, German, French, Italian, Portuguese 5 pairs · 4 sets each
Audio that alternates between English and the second language mid-stream, built from public read-speech corpora in short, long and split-long clips so the switch falls at different points.
Bangor Miami 56 recordings
Public corpus of spontaneous Spanish-English conversation among Miami residents who switch language mid-sentence. The hardest audio in this suite.
Miami, 10-hour selection 21 recordings
A ten-hour selection of the Miami Spanish-English conversation, with references produced independently by a third-party transcription service.

public, open source public, licensed private, ours

07Diarization

Speaker attribution on CALLHOME, AMI, DiPCo and NOTSOFAR, for systems that offer streaming diarization. Diarization error rate is the share of audio time given to the wrong speaker, missed, or wrongly marked as speech.

Table 14 Diarization error rate and cpWER. We lead on the DER mean and the cpWER mean. diarization error rate (DER) and cpWER · runs 2026-08-18 → 2026-09-01 · one row per model, most recent run
Diarization error rate and cpWER. We lead on the DER mean and the cpWER mean.
# System model / snapshot DER mean 4 sets CALLHOME AMI DiPCo NOTSOFAR cpWER mean CALLHOME AMI DiPCo NOTSOFAR
1 AssemblyAI Universal-3.5 Pro Realtime 17.62%13.82%19.61%25.57%11.49%32.65%19.71%30.09%38.99%41.81%
2 Azure Realtime STT 20.43%20.21%18.00%24.30%19.19%40.96%34.68%33.63%41.45%54.06%
3 Speechmatics Enhanced Realtime 22.16%14.58%19.96%28.37%25.74%37.14%24.69%23.01%42.59%58.25%
4 xAI Grok Streaming 25.68%15.59%31.90%35.28%19.96%41.81%25.77%37.95%44.72%58.82%
5 Speechmatics Standard Realtime 25.71%25.43%21.27%24.68%31.47%48.86%47.84%33.66%43.43%70.49%
6 AWS Transcribe Streaming 27.48%27.89%26.16%32.09%23.79%46.07%41.66%38.72%47.91%55.98%
7 Soniox Realtime 31.46%20.29%44.53%27.56%33.45%42.52%18.30%54.27%39.45%58.08%
8 Deepgram Nova-3 34.01%30.46%33.86%45.35%26.37%56.09%52.17%43.99%69.88%58.30%
9 Smallest Pulse 41.35%27.16%59.41%50.98%27.87%53.34%31.65%65.89%62.56%53.26%

Best Others The best value in each column is filled blue.

Data

CALLHOME 40 calls
Public CALLHOME telephone calls between friends and family: unscripted, fast, heavily overlapping conversation over a real phone line.
AMI 24 meetings
Public AMI Meeting Corpus: small groups of colleagues in real design and project meetings, in instrumented rooms.
NOTSOFAR 129 meetings
Public NOTSOFAR corpus: real multi-party meetings on a single distant microphone in ordinary meeting rooms.
DiPCo 10 sessions
Public Dinner Party Corpus: groups talking over dinner on far-field microphones, constant overlap and background clatter.

public, open source public, licensed private, ours

08Turn detection (EoT-Bench)

End-of-turn detection on LiveKit’s public EoT-Bench (English), separate from our voice-agent set. Our production model is included.

Table 15 EoT-Bench English: end-of-turn F1, precision, recall and endpoint latency. We lead on recall. EoT-Bench (LiveKit), English · runs 2026-08-31 → 2026-09-02 · fractions ×100
EoT-Bench English: end-of-turn F1, precision, recall and endpoint latency. We lead on recall.
# System model / snapshot F1 % Precision % Recall % P50 ms P90 ms Mean cutoff ms
1 OpenAI GPT Realtime 99.399.399.3157418617
2 Mistral Voxtral Mini Realtime 97.897.897.8156418676
3 Google Gemini 3.5 Transcribe Live 95.694.398.892711439
4 ElevenLabs Scribe v2 Realtime 95.193.998.0136016176
5 Cartesia Ink 2 (turns) 93.993.395.353891047
6 AssemblyAI Universal-3.5 Pro Realtime 90.386.899.8438135916
7 Azure Realtime STT 88.585.596.553672131
8 AWS Transcribe Streaming 86.782.498.849968420
9 Soniox Realtime 86.584.890.327496695
10 Deepgram Flux EN 84.883.188.5270631139
11 Deepgram Nova-3 77.375.083.532362765
12 Speechmatics Enhanced Realtime 76.874.283.51046134315
13 OpenAI GPT-4o mini Transcribe 71.064.195.555278363
14 Smallest Pulse 70.963.795.058692964
15 OpenAI GPT-4o Transcribe 70.863.796.356978065
16 Gladia Solaria 65.758.192.564896672
17 Speechmatics Standard Realtime 45.343.749.01004134114

Best Others The best value in each column is filled blue.

Data

EoT-Bench (LiveKit), English 400 clips
The English test set of LiveKit's public EoT-Bench: conversational turn-taking clips with human-marked turn boundaries.

public, open source public, licensed private, ours

09Independent evaluations

Third-party results on their own corpora and pipelines, reproduced for comparison.

Artificial Analysis — Speech to Text (Streaming)

AA-WER v2.2 (May 2026): 50% AA-AgentTalk, 25% VoxPopuli, 25% Earnings22. End of speech is detected with SileroVAD; latency runs from there to the final transcript. Price is per 1,000 minutes of audio, normalized across billing models. The leaderboard publishes no “last updated” date.

Table 16 30 of the 31 systems on the Artificial Analysis streaming leaderboard, at published ranks; our earlier-generation model is omitted. source = https://artificialanalysis.ai/speech-to-text/streaming · read 2026-09-09 · price shown for every system the source publishes one for
30 of the 31 systems on the Artificial Analysis streaming leaderboard, at published ranks; our earlier-generation model is omitted.
# System AA-WER index % t final s Price $/1k min
1 Muse Voice Transcribe (Meta) 3.06%0.163 s3.0
2 Cartesia Ink-2 (semantic) 3.36%0.431 s4.0
3 ElevenLabs Scribe v2 Realtime 3.59%0.141 s6.5
4 Qwen3 ASR Flash Realtime 3.73%0.476 s5.4
5 GPT Live Transcribe 3.92%0.812 s17.0
6 Grok STT Streaming 3.93%0.373 s3.3
7 Gemini 3.5 Transcribe Live 4.00%0.395 s9.0
8 AssemblyAI U3.5 Realtime Pro — min latency 4.02%0.191 s7.5
9 Cartesia Ink-2 (external endpoint) 4.02%0.067 s4.0
10 AssemblyAI U3.5 Realtime Pro — max accuracy 4.03%0.202 s7.5
12 Inworld STT 1 4.18%0.071 s1.4
13 Soniox v4 4.49%0.062 s2.0
14 Soniox v5 4.50%0.054 s2.0
15 Google Chirp 3 Streaming 4.80%1.276 s16.0
16 OpenAI GPT Realtime (Whisper) 4.89%0.688 s17.0
17 Smallest Pulse 4.98%0.132 s4.0
18 Voxtral Mini Realtime 5.24%0.682 s6.0
19 Azure Speech 5.25%0.625 s16.7
20 Nemotron 3 ASR (1120 ms) 5.36%0.418 s
21 Amazon Transcribe 5.79%0.620 s24.0
22 Nemotron 3 ASR (560 ms) 5.96%0.252 s
23 Deepgram Nova-3 Realtime 6.59%0.066 s4.8
24 Gradium 6.61%0.435 s12.0
25 Nemotron 3 ASR (80 ms, Together) 7.03%0.070 s
26 Deepgram Flux 7.39%0.021 s6.5
27 Nemotron 3 ASR (160 ms) 7.61%0.098 s
28 Gladia Solaria 7.82%1.523 s12.5
29 Speechmatics Realtime Enhanced 8.05%0.317 s17.5
30 Nemotron 3 ASR (80 ms) 8.38%0.073 s
31 Rev.ai 10.14%0.697 s3.3

Best Others The best value in each column is filled blue.

Table 17 The three sets in the Artificial Analysis index. The source ranks our two configurations separately; both are shown beside each set’s leading system. source = https://artificialanalysis.ai/speech-to-text/streaming · read 2026-09-09
The three sets in the Artificial Analysis index. The source ranks our two configurations separately; both are shown beside each set’s leading system.
SetOur configurationWERRankLeading system on the setIts WER
VoxPopuliU3.5 RT Pro — min latency2.06%3 of 31Gemini 3.5 Transcribe Live1.76%
VoxPopuliU3.5 RT Pro — max accuracy2.09%4 of 31Gemini 3.5 Transcribe Live1.76%
AA-AgentTalkU3.5 RT Pro — min latency3.53%9 of 31Qwen3 ASR Flash Realtime2.49%
AA-AgentTalkU3.5 RT Pro — max accuracy3.48%7 of 31Qwen3 ASR Flash Realtime2.49%
Earnings22U3.5 RT Pro — min latency6.96%11 of 31Muse Voice Transcribe4.31%
Earnings22U3.5 RT Pro — max accuracy7.08%12 of 31Muse Voice Transcribe4.31%
0.02 0.05 0.1 0.2 0.5 1 2% 4% 6% 8% 10% Time to final transcript after end of speech (s) AA-WER index (%) Muse Voice Transcribe (Meta) — 0.163 s · 3.06% Muse Voice Transcribe (Meta) 0.163 s · 3.06% Cartesia Ink-2 (semantic) — 0.431 s · 3.36% Cartesia Ink-2 (semantic) 0.431 s · 3.36% ElevenLabs Scribe v2 Realtime — 0.141 s · 3.59% 1 Qwen3 ASR Flash Realtime — 0.476 s · 3.73% Qwen3 ASR Flash Realtime 0.476 s · 3.73% GPT Live Transcribe — 0.812 s · 3.92% GPT Live Transcribe 0.812 s · 3.92% Grok STT Streaming — 0.373 s · 3.93% Grok STT Streaming 0.373 s · 3.93% Gemini 3.5 Transcribe Live — 0.395 s · 4.00% 2 U3.5 RT Pro — min latency — 0.191 s · 4.02% U3.5 RT Pro — min latency 0.191 s · 4.02% Cartesia Ink-2 (external endpoint) — 0.067 s · 4.02% Cartesia Ink-2 (external endpoint) 0.067 s · 4.02% U3.5 RT Pro — max accuracy — 0.202 s · 4.03% U3.5 RT Pro — max accuracy 0.202 s · 4.03% Inworld STT 1 — 0.071 s · 4.18% Inworld STT 1 0.071 s · 4.18% Soniox v4 — 0.062 s · 4.49% Soniox v4 0.062 s · 4.49% Soniox v5 — 0.054 s · 4.50% Soniox v5 0.054 s · 4.50% Google Chirp 3 Streaming — 1.276 s · 4.80% 3 OpenAI GPT Realtime (Whisper) — 0.688 s · 4.89% 4 Smallest Pulse — 0.132 s · 4.98% Smallest Pulse 0.132 s · 4.98% Voxtral Mini Realtime — 0.682 s · 5.24% Voxtral Mini Realtime 0.682 s · 5.24% Azure Speech — 0.625 s · 5.25% Azure Speech 0.625 s · 5.25% Nemotron 3 ASR (1120 ms) — 0.418 s · 5.36% Nemotron 3 ASR (1120 ms) 0.418 s · 5.36% Amazon Transcribe — 0.62 s · 5.79% Amazon Transcribe 0.62 s · 5.79% Nemotron 3 ASR (560 ms) — 0.252 s · 5.96% 5 Deepgram Nova-3 Realtime — 0.066 s · 6.59% Deepgram Nova-3 Realtime 0.066 s · 6.59% Gradium — 0.435 s · 6.61% Gradium 0.435 s · 6.61% Nemotron 3 ASR (80 ms, Together) — 0.07 s · 7.03% Nemotron 3 ASR (80 ms, Together) 0.07 s · 7.03% Deepgram Flux — 0.021 s · 7.39% Deepgram Flux 0.021 s · 7.39% Nemotron 3 ASR (160 ms) — 0.098 s · 7.61% Nemotron 3 ASR (160 ms) 0.098 s · 7.61% Gladia Solaria — 1.523 s · 7.82% Gladia Solaria 1.523 s · 7.82% Speechmatics Realtime Enhanced — 0.317 s · 8.05% Speechmatics Realtime Enhanced 0.317 s · 8.05% Nemotron 3 ASR (80 ms) — 0.073 s · 8.38% Nemotron 3 ASR (80 ms) 0.073 s · 8.38% Rev.ai — 0.697 s · 10.14% Rev.ai 0.697 s · 10.14%
  1. 1 ElevenLabs Scribe v2 Realtime · 0.141 s · 3.59%
  2. 2 Gemini 3.5 Transcribe Live · 0.395 s · 4.00%
  3. 3 Google Chirp 3 Streaming · 1.276 s · 4.80%
  4. 4 OpenAI GPT Realtime (Whisper) · 0.688 s · 4.89%
  5. 5 Nemotron 3 ASR (560 ms) · 0.252 s · 5.96%
Figure 9 AA-WER against time to final transcript. The dashed line is the accuracy–latency frontier. source = https://artificialanalysis.ai/speech-to-text/streaming

Coval — Live STT Leaderboard

Rolling 24-hour window (the source’s default view), read 2026-09-09 with 28 systems ranked. WER is pooled over seven conditions (clean, accent, clipping, far-field, noise gap, phone codec, reverb), 44–45 clips each and 176 for clean. The stt-v3 references are model-generated, not human-verified. Coval sells voice-agent evaluation tooling.

Table 18 Selected rows of 28. We rank 1st of 28 on pooled WER. source = https://benchmarks.coval.ai/stt · read 2026-09-09 · rolling 24 h window
Selected rows of 28. We rank 1st of 28 on pooled WER.
System model / snapshot WER pooled %
AssemblyAI universal-3.5-pro rank 1 of 28 2.91%
Baseten qwen3-asr-1.7b rank 2 of 28 3.19%
Reson8 realtime rank 3 of 28 3.34%
Google Chirp 3 rank 4 of 28 4.15%
Inworld rank 5 of 28 4.18%
ElevenLabs Scribe v2 Realtime rank 16 of 28 5.69%
Soniox stt-rt-v5 rank 19 of 28 6.21%
Deepgram Flux (general-en) rank 21 of 28 6.67%

Best Others The best value in each column is filled blue.

Figure 10 Our WER by recording condition. The dashed line is the pooled value from Table 18. The axis is clipped at 4%; reverb exceeds it. source = https://benchmarks.coval.ai/stt
0 1 2 3 4 accent: 2.07% accent 2.07% phone codec: 2.33% phone codec 2.33% clean: 2.46% clean 2.46% noise gap: 2.53% noise gap 2.53% far-field: 2.61% far-field 2.61% clipping: 2.90% clipping 2.90% reverb: 13.64% reverb › 13.64% pooled 2.91%

Pipecat / Daily — STT benchmark

1,000 English samples. Semantic WER is pooled over all reference words; time to final is the median and P95 from end of speech to the final segment. Read from the repository README (last changed 2026-08-27). Vendors without a model name in the source are listed by vendor only; previous-generation AssemblyAI models are not shown.

Table 19 Shown as published; our row is the production model named in the source. source = https://github.com/pipecat-ai/stt-benchmark · read 2026-09-09
Shown as published; our row is the production model named in the source.
# System Semantic WER % TTF median ms TTF p95 ms
1 Speechmatics 1.07%495676
2 Azure 1.18%10161345
3 AssemblyAI universal-3.5-pro 1.22%282354
4 Cartesia ink-2 1.25%299328
5 Soniox stt-rt-v5 1.27%260305
6 Soniox stt-rt-v4 1.29%249281
7 Deepgram nova-3-general 1.62%247298
8 AWS Transcribe 1.75%11361527
9 NVIDIA Nemotron 3.0 ASR (en) 1.95%221238
10 Google gemini-3.5-transcribe-live 2.24%458532
11 Smallest AI pulse 2.37%398533
12 OpenAI gpt-realtime-whisper 2.73%740878
13 Google latest-long 2.85%8781155
14 OpenAI gpt-4o-transcribe 3.06%637965
15 ElevenLabs scribe_v2_realtime 3.12%281348
16 Gradium 3.71%570596
17 Cartesia ink-whisper 4.36%266364
18 NVIDIA Nemotron 3.5 ASR (multilingual) 4.58%236253
19 Mistral voxtral-mini-transcribe-realtime-2602 4.97%525973

Best Others The best value in each column is filled blue.

10How we measure

How the results were produced: the harness, the metrics, why public leaderboards can rank the same systems differently, and the limits of this report.

How we run every system

Protocol. Audio is streamed in real time to each system’s public API, default settings, language fixed. One normalizer is applied before scoring. Each streaming model a vendor ships has its own row.

Complete runs only. An average is shown only when the system was scored on every dataset in it; partial rows keep their cells but are not ranked.

n/d, n/s. A 0.00 error rate means no data. WER or CER of 60% or more means the language is unsupported; the cell is excluded from means.

Excluded. Unreleased models, and results on a superseded version of the voice-agent test set.

Systems

Every system with a public streaming API and a current run; a dot marks each suite it was scored on.

Table 20 Systems evaluated and the suites each has a current run on. 6 suites · runs 2026-08-27 → 2026-09-08
Systems evaluated and the suites each has a current run on.
VendorModelEnglish / CIVoice agentsCode-switch.DiarizationMultilingualTurn detect.Runs
AssemblyAIUniversal-3.5 Pro Realtime2026-08-18 → 2026-09-02
AWSTranscribe Streaming2026-08-19 → 2026-08-31
AzureAzure Realtime STT·2026-08-19 → 2026-08-31
CartesiaInk 2 (turns)·····2026-08-31
CartesiaInk-Whisper···2026-08-19 → 2026-08-30
CartesiaInk-Whisper EN·····2026-08-20
DeepgramFlux·····2026-08-25
DeepgramFlux EN···2026-08-19 → 2026-08-31
DeepgramFlux Multi·····2026-08-27
DeepgramNova-3·2026-08-18 → 2026-08-31
DeepgramNova-3 EN·····2026-08-20
DeepgramNova-3 Multi·····2026-08-20
ElevenLabsScribe v2 Realtime·2026-08-19 → 2026-09-01
GladiaSolaria··2026-08-19 → 2026-08-31
GoogleGemini 3.5 Transcribe Live·2026-08-31 → 2026-09-06
MistralVoxtral Mini Realtime·2026-08-19 → 2026-08-31
OpenAIGPT Realtime·2026-08-19 → 2026-08-31
OpenAIGPT-4o Transcribe·2026-08-19 → 2026-08-31
OpenAIGPT-4o mini Transcribe·2026-08-19 → 2026-08-31
Reson8Reson8·····2026-09-01
SmallestPulse2026-08-28 → 2026-08-31
SonioxSoniox Realtime·2026-08-20 → 2026-08-31
SpeechmaticsEnhanced Realtime2026-08-19 → 2026-08-31
SpeechmaticsStandard Realtime2026-08-19 → 2026-08-31
xAIGrok Streaming···2026-08-19 → 2026-09-08

Test data

The audio each benchmark uses and the metric it scores. Each results section also lists its own datasets.

Table 21 Test data by benchmark. Files are scored recordings or clips; multilingual, code-switching and judge-scored sets are scored per language or label, so no single count applies. Public corpora are named, licensed sets described.
Test data by benchmark. Files are scored recordings or clips; multilingual, code-switching and judge-scored sets are scored per language or label, so no single count applies. Public corpora are named, licensed sets described.
SuiteAudioFilesScored on
Voice agentsScripted caller scenarios voiced by real speakers, recorded in three noise environments, three utterance-length bands and 12 speaker groups12,460Entity error, WER, end-of-turn detection
HealthcarePriMock57 primary-care consultations; licensed clinician dictation; respiratory-clinic recordings479Medical entity error
English short-formCommon Voice; LibriSpeech test-clean and test-other21,588WER
English accentedBritish-, French-, Spanish- and Indian-accented English9,315WER
English long-formTED-LIUM talks; Rev16 podcasts; CALLHOME telephone calls91WER
Judge-scored entitiesShort-form English audio, with entities graded per label by an LLM judgeEntity error (judge)
Non-speechHold music, line noise and silence with no speech2,950Response rate
MultilingualFLEURS and Common Voice test sets in nine languagesWER (character error rate for Japanese, Chinese)
Code-switchingFive English↔X pairs derived from read speech, plus Bangor Miami bilingual conversationWER
DiarizationCALLHOME calls; AMI and NOTSOFAR meetings; DiPCo dinner parties203DER, cpWER
Turn detectionEoT-Bench (LiveKit), English400End-of-turn F1, endpoint latency

Each test set is marked by how its audio can be obtained: ○ public and open source; ◆ public but licensed or purchasable; ● private, recorded or annotated by us.

How we score

Each metric and how it is computed. The same normalizer and scoring code apply to every system.

WER Word error rate after normalization; lower is better. Character error rate for Japanese, Chinese and Korean.

NEER Normalized entity error rate: the share of entities (names, addresses, numbers, emails, dates, codes, medical and technical terms) transcribed incorrectly.

MER Medical entity error rate over medication and disease terms.

Judge-scored entity error Entity error rate as graded by an LLM judge.

End-of-turn F1 / P / R Whether the endpoint fires exactly once per real turn. Latency is the time (ms) from true end of speech to the endpoint event.

DER / cpWER Diarization error rate, and concatenated minimum-permutation WER.

Non-speech response rate Share of non-speech clips for which the system produces any text.

Streaming harness. Audio is chunked and streamed in real time to each vendor’s endpoint through its documented SDK or WebSocket API, default settings, language fixed. Final transcripts are scored, not partials.

Normalization. One normalizer (case, punctuation, numbers, abbreviations) is applied to references and to every system’s output. Entity metrics are computed on normalized text and, where reported, on formatted text.

Entity error. Entity spans are located in the reference; an entity is an error if any token of its span is wrong in the aligned transcript. The judge-scored variant asks an LLM whether the transcript names the same real-world entity.

Endpointing. An endpoint is a true positive if it falls inside the tolerance window after the reference end of turn. Latency is its offset from that end of turn.

Selection rule. Every row is the model’s most recent production run. The overview grid keeps each vendor’s best model per column and names it on hover.

Provenance. Each figure carries its metric and run-date span. Public test sets are named; licensed sets are described.

Why public leaderboards and ours differ

We build for voice agents, clinical documentation, meetings, telephony, and multilingual and code-switched speech, so the test sets cover that range of speakers, accents, languages and acoustic conditions and score what matters in each. Public leaderboards usually ask a narrower question, typically one word-error-rate figure on a few corpora, so a model can rank differently on the two.

Table 22 What each benchmark measures: this report and the public leaderboards reproduced in §9.
What each benchmark measures: this report and the public leaderboards reproduced in §9.
BenchmarkAudioScoresRun byReferences
This reportVoice-agent calls in three noise environments, clinical consultations and dictation, accented and long-form English, nine languages, code-switched speech, meetings and telephone calls, non-speech audioEntity error (by class), WER, DER, end-of-turn F1 and latency, non-speech response rateAssemblyAI; every vendor on default settingsHuman-verified
Artificial AnalysisAA-AgentTalk (50%), VoxPopuli parliamentary speech (25%), Earnings22 calls (25%)One WER index over the three sets; time to final transcriptArtificial AnalysisPublished with the sets
CovalShort clips under seven recording conditions (clean, accent, far-field, phone codec, reverb…), rolling 24-hour windowPooled WER; time to first token; time to final segmentCovalModel-generated
Pipecat / Daily1,000 English samplesSemantic WER; median and P95 time to finalDailyPublished with the set
EoT-Bench (LiveKit)Scripted turns in 14 languagesFalse-cutoff rate at 300 and 600 msLiveKitScripted

The Artificial Analysis index is one WER over three sets, half of it AA-AgentTalk, with no entity, noise, medical, multilingual, code-switching or diarization component; its latency runs from a voice-activity detector’s end of speech to the final transcript, a different quantity from the endpointing latency in §2. Coval scores against model-generated references. Neither substitutes for the other or for this report.

Caveats

We ran it. AssemblyAI produced the results in §1–§8. Competitors ran on default settings and may do better tuned. §9 has independent results.

Snapshots differ. Runs span 2026-08-18 to 2026-09-08 (Table 20). A vendor that updated its model within that window may appear on the older version.

Coverage is uneven. Not every system ran every set; means are withheld rather than computed on partial data.

Reference quality. TED-LIUM references omit sponsor read-outs. Some independent leaderboards (Coval) use model-generated references.

Latency is measured one way. Endpointing latency here runs from end of speech to the endpoint event. Independent leaderboards measure time to final transcript or to first token, which are not comparable.

Benchmarks

Run your own benchmark

Testing on your own audio? Want to do it right? Read our benchmarking guide in the docs.