New Universal-3.6 Pro Realtime is now available Learn more
Benchmarks

AssemblyAI Universal-3.6 Pro Realtime benchmarks

Universal-3.6 Pro Realtime measured against every major streaming vendor on the same audio.

Technical report Realtime speech-to-text data as of 2026-09-16 v2026.09

Realtime speech recognition, measured against the field

Universal-3.6 Pro Realtime and 26 competitor models on 5 evaluation suites: same audio, same scoring pipeline.

Universal-3.6 Pro Realtime has the lowest entity error rate on the voice-agent suite: 14.41% against 18.19% for the nearest competing system, a 21% lead. It also leads on code-switched speech.

Every score is in the tables. §9 reproduces independent evaluations; §10 describes the method.

At a glance

Voice-agent entity error
14.41%
1st of 22
next best Soniox Realtime 18.19%
Code-switching (5 pairs)
7.20%
1st of 23
next best AWS Transcribe 9.13%
Voice-agent WER, heavy noise
6.13%
1st of 22
next best Soniox Realtime 7.68%
English short-form WER
3.44%
1st of 22
next best Reson8 3.70%
Voice-agent WER, Indian-accented
6.64%
1st of 22
next best Soniox Realtime 7.78%

01Overview

Each cell is the vendor’s best model on that metric, coloured by rank within the column (solid blue is first). Values are in the tables of §2–§8.

Rank across 9 headline metrics, vendors ordered by mean rank.
Vendor (best model per cell) VA entity err. NEER % Entity (judge) % Medical MER % Diarization DER % Non-speech resp. % Code-switching WER % Multi-lingual FLEURS % EN short-form WER % EN long-form WER % meanrank
AssemblyAI 14.418.14.216.00.37.28.43.48.8 3.0
Google 20.119.78.6—0.210.26.45.09.4 5.3
Meta 22.220.54.114.50.313.310.06.27.4 5.9
Reson8 23.724.414.316.70.710.75.83.77.7 6.0
Speechmatics 20.619.84.722.20.238.211.14.78.1 6.6
Soniox 18.221.54.631.80.611.37.97.68.2 6.7
AWS 19.328.110.227.51.39.18.93.98.7 7.2
ElevenLabs 18.520.47.0—0.213.09.35.117.1 7.4
Zoom 18.224.910.5—0.332.56.23.89.3 7.9
Mistral 23.120.812.3—0.018.59.56.618.9 10.0
Azure 29.733.114.420.40.214.612.75.59.2 10.2
Deepgram 26.125.913.234.00.217.112.76.59.6 10.4
xAI 24.026.115.725.71.220.614.85.89.4 11.6
OpenAI 38.120.217.2—0.133.312.38.015.9 12.5
Smallest 29.433.221.441.41.414.611.95.410.9 12.6
Gladia 30.624.917.0—4.917.116.27.010.7 14.1
Cartesia 40.127.821.1—3.150.518.77.710.6 16.1

Best Worst Cells are coloured by rank within each column (solid blue is best, white is worst).

Figure 1 Rank across 9 headline metrics, vendors ordered by mean rank. source = the tables below; each column’s provenance is given with its table

Position in the field

  • Voice-agent entity error 14.41% · 1st of 22 systems · nearest competitor Soniox Realtime 18.19% −20.8%
  • Judge-scored entity error 18.93% · 2nd of 22 systems · nearest competitor Google Gemini 3.5 Live 19.68% −3.8%
  • Medical entity error 4.24% · 2nd of 22 systems · nearest competitor Meta Muse Voice Transcribe 4.08% +3.9%
  • Diarization DER 16.68% · 3rd of 12 systems · nearest competitor Meta Muse Voice Transcribe 14.51% +15.0%
  • Non-speech response rate 0.27% · 9th of 22 systems · nearest competitor Mistral Voxtral Mini 0.03% +0.2 pt
  • Code-switching (5 pairs) 7.20% · 1st of 23 systems · nearest competitor AWS Transcribe 9.13% −21.1%
  • Multilingual (FLEURS) 8.44% over 18 lang. · 5th of 22 systems · nearest competitor Reson8 5.77% over 7 lang. +46.3%
  • English short-form WER 3.44% · 1st of 22 systems · nearest competitor Reson8 3.70% −7.0%
  • Long-form English WER 8.82% · 6th of 22 systems · nearest competitor Meta Muse Voice Transcribe 7.37% +19.7%

We lead We trail The margin is blue where we lead, grey where we trail.

One row per headline metric from the grid above: our value, our rank among systems with a complete run, and the margin to the runner-up where we lead or to the leader where we do not. Margins are relative for error rates and absolute for rates below one percent.

02Voice agents

Conversational audio as a voice agent hears it, across 22 systems, three noise environments, 12 speaker groups and three utterance lengths. The headline metric, entity error (NEER), is how often the names, addresses and numbers an agent acts on come out wrong.

Table 1 Entity error rate (%) by entity class. Colour is rank within the column. We lead in 9 of 11 columns. normalized entity error rate (NEER) · runs 2026-08-27 → 2026-09-22 · one row per model, most recent run · corpus value = mean of the three natural-noise environments
Entity error rate (%) by entity class. Colour is rank within the column. We lead in 9 of 11 columns.
System Overall Names Places Addresses Phone Email Dates Numbers Codes/IDs Medical Technical
AssemblyAI Universal-3.6 Pro 14.410.96.029.02.417.218.612.910.016.120.9
AssemblyAI Universal-3.5 Pro 14.811.27.327.92.617.119.113.210.217.321.7
Soniox Realtime 18.216.67.727.92.722.730.819.78.219.126.5
Zoom Scribe (Live) 18.216.57.030.09.622.420.914.913.920.026.4
ElevenLabs Scribe v2 18.514.88.928.63.421.226.517.212.024.727.2
AWS Transcribe 19.321.37.326.83.121.927.818.813.019.033.7
Google Gemini 3.5 Live 20.115.89.332.47.323.930.417.316.920.527.4
Speechmatics Enhanced 20.623.86.031.63.224.231.117.119.417.431.8
Meta Muse Voice Transcribe 22.223.98.438.510.020.020.017.427.825.630.5
Mistral Voxtral Mini 23.118.310.338.97.727.528.617.316.131.634.3
Reson8 23.718.68.930.811.931.826.618.720.938.031.3
xAI Grok Streaming 24.017.911.840.212.430.019.918.120.035.734.3
Deepgram Nova-3 26.124.311.030.44.534.443.716.827.629.139.4
Smallest Pulse 29.427.515.441.812.139.524.919.533.238.542.0
Azure 29.732.417.953.68.722.230.523.535.432.440.5
Deepgram Flux EN 30.129.013.241.011.540.722.821.046.132.343.9
Deepgram Flux Multi 30.432.114.241.79.336.222.322.346.037.143.1
Gladia Solaria 30.621.313.539.620.642.431.619.237.440.540.0
Speechmatics Standard 30.630.69.341.66.857.132.520.143.328.836.2
OpenAI GPT-4o Transcribe 38.127.817.951.742.850.333.230.547.737.341.8
OpenAI GPT-4o mini Transcribe 38.929.319.653.340.652.135.628.346.042.541.9
Cartesia Ink-Whisper 40.125.114.656.958.947.133.625.352.645.142.1

Best Worst Cells are coloured by rank within each column (solid blue is best, white is worst).

Background noise and utterance length

0 10 20 30 AssemblyAI Universal-3.6 Pro · No background noise: 12.4% AssemblyAI Universal-3.6 Pro · Natural low noise: 13.3% AssemblyAI Universal-3.6 Pro · Natural heavy noise: 13.6% AssemblyAI Universal-3.6 Pro AssemblyAI Universal-3.5 Pro · No background noise: 12.5% AssemblyAI Universal-3.5 Pro · Natural low noise: 13.4% AssemblyAI Universal-3.5 Pro · Natural heavy noise: 14.1% AssemblyAI Universal-3.5 Pro Zoom Scribe (Live) · No background noise: 16.1% Zoom Scribe (Live) · Natural low noise: 17.0% Zoom Scribe (Live) · Natural heavy noise: 18.0% Zoom Scribe (Live) ElevenLabs Scribe v2 · No background noise: 16.1% ElevenLabs Scribe v2 · Natural low noise: 17.6% ElevenLabs Scribe v2 · Natural heavy noise: 18.6% ElevenLabs Scribe v2 Soniox Realtime · No background noise: 17.9% Soniox Realtime · Natural low noise: 18.7% Soniox Realtime · Natural heavy noise: 18.8% Soniox Realtime Google Gemini 3.5 Live · No background noise: 18.1% Google Gemini 3.5 Live · Natural low noise: 19.4% Google Gemini 3.5 Live · Natural heavy noise: 20.7% Google Gemini 3.5 Live xAI Grok Streaming · No background noise: 19.0% xAI Grok Streaming · Natural low noise: 20.2% xAI Grok Streaming · Natural heavy noise: 21.9% xAI Grok Streaming AWS Transcribe · No background noise: 19.7% AWS Transcribe · Natural low noise: 20.5% AWS Transcribe · Natural heavy noise: 21.5% AWS Transcribe Mistral Voxtral Mini · No background noise: 19.8% Mistral Voxtral Mini · Natural low noise: 20.9% Mistral Voxtral Mini · Natural heavy noise: 22.3% Mistral Voxtral Mini Reson8 · No background noise: 20.4% Reson8 · Natural low noise: 22.3% Reson8 · Natural heavy noise: 22.6% Reson8 Meta Muse Voice Transcribe · No background noise: 21.5% Meta Muse Voice Transcribe · Natural low noise: 22.3% Meta Muse Voice Transcribe · Natural heavy noise: 23.2% Meta Muse Voice Transcribe Speechmatics Enhanced · No background noise: 22.1% Speechmatics Enhanced · Natural low noise: 23.0% Speechmatics Enhanced · Natural heavy noise: 23.4% Speechmatics Enhanced Gladia Solaria · No background noise: 24.6% Gladia Solaria · Natural low noise: 27.6% Gladia Solaria · Natural heavy noise: 29.1% Gladia Solaria Smallest Pulse · No background noise: 26.0% Smallest Pulse · Natural low noise: 28.6% Smallest Pulse · Natural heavy noise: 27.6% Smallest Pulse Deepgram Nova-3 · No background noise: 26.8% Deepgram Nova-3 · Natural low noise: 27.7% Deepgram Nova-3 · Natural heavy noise: 29.2% Deepgram Nova-3 Deepgram Flux EN · No background noise: 27.5% Deepgram Flux EN · Natural low noise: 28.6% Deepgram Flux EN · Natural heavy noise: 30.0% Deepgram Flux EN Deepgram Flux Multi · No background noise: 28.3% Deepgram Flux Multi · Natural low noise: 29.8% Deepgram Flux Multi · Natural heavy noise: 31.6% Deepgram Flux Multi Azure · No background noise: 29.1% Azure · Natural low noise: 30.3% Azure · Natural heavy noise: 31.6% Azure Speechmatics Standard · No background noise: 29.5% Speechmatics Standard · Natural low noise: 31.4% Speechmatics Standard · Natural heavy noise: 32.0% Speechmatics Standard OpenAI GPT-4o Transcribe · No background noise: 31.2% OpenAI GPT-4o Transcribe · Natural low noise: 33.6% OpenAI GPT-4o Transcribe · Natural heavy noise: 34.5% OpenAI GPT-4o Transcribe Cartesia Ink-Whisper · No background noise: 31.3% Cartesia Ink-Whisper · Natural low noise: 33.5% Cartesia Ink-Whisper · Natural heavy noise: 35.4% Cartesia Ink-Whisper OpenAI GPT-4o mini Transcribe · No background noise: 32.1% OpenAI GPT-4o mini Transcribe · Natural low noise: 34.8% OpenAI GPT-4o mini Transcribe · Natural heavy noise: 35.3% OpenAI GPT-4o mini Transcribe
No background noise Natural low noise Natural heavy noise
Figure 2 Entity error rate (%) by background-noise condition, sorted by the no-noise value. We lead in all 3 conditions. normalized entity error rate (NEER) · runs 2026-08-27 → 2026-09-22 · one row per model, most recent run

Listen to the test data

One caller turn per noise environment, quietest first, each with an entity an agent acts on: a name, an account number, an e-mail login, a medication. Words that differ from the human reference are marked.

1 of 8
Voice agentsno noise4 s · AssemblyAI voice-agent corpus — scripted caller scenarios voiced by real speakers

Four seconds, no noise. A short request where the other systems substitute unrelated words.

Reference

My cable modem keeps rebooting every 30-40 minutes, can you help?

AssemblyAI

My cable modem keeps rebooting every 30 to 40 minutes. Can you help?

ElevenLabs Scribe v2 Realtime

My key won't let them keep changing every 30 to 40 minutes. Can you help?

Soniox Realtime

Mark Kilmer and keeps removing every 30 to 40 minutes. Can you help?

0 10 20 30 40 AssemblyAI Universal-3.6 Pro · Short utterances: 12.6% AssemblyAI Universal-3.6 Pro · Medium: 12.5% AssemblyAI Universal-3.6 Pro · Long: 15.1% AssemblyAI Universal-3.6 Pro AssemblyAI Universal-3.5 Pro · Short utterances: 12.7% AssemblyAI Universal-3.5 Pro · Medium: 12.7% AssemblyAI Universal-3.5 Pro · Long: 15.6% AssemblyAI Universal-3.5 Pro ElevenLabs Scribe v2 · Short utterances: 16.8% ElevenLabs Scribe v2 · Medium: 16.9% ElevenLabs Scribe v2 · Long: 19.8% ElevenLabs Scribe v2 Zoom Scribe (Live) · Short utterances: 17.4% Zoom Scribe (Live) · Medium: 16.2% Zoom Scribe (Live) · Long: 18.2% Zoom Scribe (Live) Google Gemini 3.5 Live · Short utterances: 17.9% Google Gemini 3.5 Live · Medium: 18.8% Google Gemini 3.5 Live · Long: 23.3% Google Gemini 3.5 Live xAI Grok Streaming · Short utterances: 19.2% xAI Grok Streaming · Medium: 19.6% xAI Grok Streaming · Long: 23.9% xAI Grok Streaming Soniox Realtime · Short utterances: 19.9% Soniox Realtime · Medium: 17.4% Soniox Realtime · Long: 18.3% Soniox Realtime Reson8 · Short utterances: 21.1% Reson8 · Medium: 21.3% Reson8 · Long: 23.7% Reson8 AWS Transcribe · Short utterances: 21.2% AWS Transcribe · Medium: 20.0% AWS Transcribe · Long: 20.7% AWS Transcribe Mistral Voxtral Mini · Short utterances: 21.5% Mistral Voxtral Mini · Medium: 19.7% Mistral Voxtral Mini · Long: 22.9% Mistral Voxtral Mini Meta Muse Voice Transcribe · Short utterances: 23.0% Meta Muse Voice Transcribe · Medium: 21.2% Meta Muse Voice Transcribe · Long: 23.6% Meta Muse Voice Transcribe Speechmatics Enhanced · Short utterances: 23.1% Speechmatics Enhanced · Medium: 22.1% Speechmatics Enhanced · Long: 23.8% Speechmatics Enhanced Smallest Pulse · Short utterances: 26.1% Smallest Pulse · Medium: 26.8% Smallest Pulse · Long: 31.0% Smallest Pulse Gladia Solaria · Short utterances: 27.1% Gladia Solaria · Medium: 26.5% Gladia Solaria · Long: 28.6% Gladia Solaria Azure · Short utterances: 27.5% Azure · Medium: 29.4% Azure · Long: 37.0% Azure Deepgram Nova-3 · Short utterances: 28.2% Deepgram Nova-3 · Medium: 27.5% Deepgram Nova-3 · Long: 28.2% Deepgram Nova-3 Deepgram Flux EN · Short utterances: 28.8% Deepgram Flux EN · Medium: 27.9% Deepgram Flux EN · Long: 30.1% Deepgram Flux EN Deepgram Flux Multi · Short utterances: 30.2% Deepgram Flux Multi · Medium: 29.2% Deepgram Flux Multi · Long: 30.8% Deepgram Flux Multi OpenAI GPT-4o Transcribe · Short utterances: 30.5% OpenAI GPT-4o Transcribe · Medium: 32.8% OpenAI GPT-4o Transcribe · Long: 38.0% OpenAI GPT-4o Transcribe Speechmatics Standard · Short utterances: 30.7% Speechmatics Standard · Medium: 30.5% Speechmatics Standard · Long: 32.4% Speechmatics Standard OpenAI GPT-4o mini Transcribe · Short utterances: 31.7% OpenAI GPT-4o mini Transcribe · Medium: 33.6% OpenAI GPT-4o mini Transcribe · Long: 39.1% OpenAI GPT-4o mini Transcribe Cartesia Ink-Whisper · Short utterances: 32.4% Cartesia Ink-Whisper · Medium: 33.0% Cartesia Ink-Whisper · Long: 35.8% Cartesia Ink-Whisper
Short utterances Medium Long
Figure 3 Entity error rate (%) by utterance length. We lead in all 3 length bands. normalized entity error rate (NEER) · runs 2026-08-27 → 2026-09-22 · one row per model, most recent run

Speaker groups

Table 2 Entity error rate (%) by speaker group. Colour is rank within the column. We lead in all 12 groups. normalized entity error rate (NEER) · runs 2026-08-27 → 2026-09-22 · one row per model, most recent run
Entity error rate (%) by speaker group. Colour is rank within the column. We lead in all 12 groups.
System US Northeast US South US Gen. Am. US AAVE US Asian-Am. US other UK Canada Native other Spanish-acc. Indian-acc. Non-nat. other meanrank
AssemblyAI Universal-3.6 Pro 10.111.511.412.412.013.914.010.515.216.119.118.9 1.0
AssemblyAI Universal-3.5 Pro 10.111.611.612.812.214.514.211.015.216.620.319.1 2.0
Zoom Scribe (Live) 13.815.815.015.315.820.418.713.519.120.523.924.1 3.4
ElevenLabs Scribe v2 14.615.815.016.616.119.419.513.719.822.923.825.4 4.2
Soniox Realtime 16.316.516.518.118.920.719.315.820.522.423.824.8 5.3
Google Gemini 3.5 Live 17.218.617.820.716.822.421.515.620.623.621.225.8 6.6
AWS Transcribe 17.718.419.220.719.822.420.118.021.623.824.427.4 7.7
xAI Grok Streaming 15.019.418.320.617.723.223.015.424.422.427.628.2 8.0
Mistral Voxtral Mini 18.019.018.320.419.622.420.615.523.628.728.031.1 8.8
Reson8 17.720.519.219.720.025.621.616.723.427.728.532.0 9.8
Meta Muse Voice Transcribe 19.520.320.621.521.122.722.719.224.128.026.629.6 10.5
Speechmatics Enhanced 21.421.021.222.022.124.822.619.823.827.026.629.6 10.8
Gladia Solaria 22.025.224.627.327.329.628.020.530.333.532.737.6 14.0
Deepgram Nova-3 24.426.426.227.027.228.428.624.330.131.032.935.1 14.3
Smallest Pulse 21.625.524.726.029.430.126.521.530.234.631.040.1 14.7
Deepgram Flux EN 24.426.526.027.527.130.432.225.732.333.734.937.6 15.8
Azure 26.629.028.429.028.932.131.426.031.936.435.837.8 17.5
Deepgram Flux Multi 25.327.227.129.227.732.633.026.333.336.136.339.8 17.8
Speechmatics Standard 26.728.928.431.632.033.431.528.233.537.435.739.7 18.8
OpenAI GPT-4o Transcribe 28.932.730.031.430.735.734.925.136.342.339.844.2 19.9
Cartesia Ink-Whisper 27.932.230.531.932.735.734.025.836.841.640.944.1 20.4
OpenAI GPT-4o mini Transcribe 29.433.030.532.332.437.035.327.537.542.941.646.9 21.8

Best Worst Cells are coloured by rank within each column (solid blue is best, white is worst).

Ablations on our own model

wordboost OFF wordboost ON
Figure 4 Word boost (custom vocabulary) on versus off, same model and corpus. Change in entity error with word boost on: names −60.9%, technical terms −72.8%. normalized entity error rate (NEER) + WER, wordboost OFF vs ON · runs 2026-09-22 · one row per model, most recent run
agent context OFF agent context ON
Figure 5 Agent context on versus off: the agent’s previous turn is passed to the recognizer. normalized entity error rate (NEER) + WER, same model with and without agent context · runs 2026-08-27 → 2026-08-28 · one row per model, most recent run

Word error rate

Table 3 WER by noise condition. normalized word error rate · runs 2026-08-27 → 2026-09-22 · one row per model, most recent run
WER by noise condition.
# System model / snapshot Mean 3 conditions No noise Low noise Heavy noise
1 AssemblyAI Universal-3.6 Pro Realtime 5.19%4.37%5.08%6.13%
2 AssemblyAI Universal-3.5 Pro Realtime 5.72%4.61%5.47%7.08%
3 Soniox Realtime 6.83%5.98%6.84%7.68%
4 Meta Muse Voice Transcribe 7.48%6.92%7.50%8.02%
5 ElevenLabs Scribe v2 Realtime 7.78%6.18%7.53%9.64%
6 Speechmatics Enhanced Realtime 8.14%6.96%7.96%9.49%
7 Google Gemini 3.5 Transcribe Live 8.48%7.44%8.49%9.52%
8 AWS Transcribe Streaming 8.58%7.27%8.27%10.19%
9 Deepgram Nova-3 8.64%7.53%8.45%9.93%
10 Zoom Scribe (Live) 9.36%8.04%9.29%10.74%
11 Reson8 Reson8 9.41%8.21%9.38%10.64%
12 Mistral Voxtral Mini Realtime 10.43%9.25%10.45%11.59%
13 Smallest Pulse 10.74%9.85%10.94%11.43%
14 Azure Realtime STT 10.86%9.78%10.65%12.15%
15 Speechmatics Standard Realtime 12.05%10.47%11.77%13.91%
16 Deepgram Flux EN 13.50%12.29%13.36%14.85%
17 Deepgram Flux Multi 14.26%12.85%14.06%15.87%
18 Cartesia Ink-Whisper 15.17%13.71%15.00%16.81%
19 Gladia Solaria 16.25%12.99%16.04%19.71%
20 OpenAI GPT-4o Transcribe 18.37%17.05%18.44%19.63%
21 OpenAI GPT-4o mini Transcribe 19.55%18.12%19.77%20.76%
22 xAI Grok Streaming 20.28%18.45%19.89%22.50%

Best Others The best value in each column is filled blue.

Data

●Caller corpus 12,460 clips
All entity classes pooled into one number: how often an agent-critical value (a name, address, number or identifier) comes back wrong.
●Noise environments
Three levels of real noise, recorded rather than mixed in. None: room and handset noise only. Low: incidental noise from a home, office or parked car. Heavy: noise that competes with the voice.
●Utterance length
Short turns (confirmations, single values, one-word answers), medium and long. Short turns give the model almost no context.
●Speaker groups
Fourteen groups: US Northeast, South, General American, AAVE, Asian-American and other; UK; Canada; native speakers overall; Spanish- and Indian-accented and other non-native speakers.
●Entity classes
Names (including letter-by-letter spellings), places, street addresses, phone numbers, email addresses, dates and times, numbers and currency, codes and identifiers, medical and technical terms.

public, open source public, licensed private, ours

03Healthcare

Medical entity error is measured on primary-care consultations (PriMock57), licensed clinical dictation and respiratory-clinic recordings.

Overall

Table 4 Medical entity error rate. We lead on primary care. normalized medical entity error rate (MER) · runs 2026-08-18 → 2026-09-22 · one row per model, most recent run
Medical entity error rate. We lead on primary care.
# System model / snapshot Mean 3 sets Primary care PriMock57 Dictation Respiratory
1 Meta Muse Voice Transcribe 4.08%7.18%3.09%1.97%
2 AssemblyAI Universal-3.6 Pro Realtime 4.24%6.23%4.20%2.28%
3 Soniox Realtime 4.58%7.18%4.04%2.51%
4 Speechmatics Enhanced Realtime 4.66%8.13%3.65%2.21%
5 AssemblyAI Universal-3.5 Pro Realtime 4.81%7.07%5.08%2.28%
6 ElevenLabs Scribe v2 Realtime 7.00%10.89%8.01%2.09%
7 Google Gemini 3.5 Transcribe Live 8.57%15.45%7.69%2.56%
8 AWS Transcribe Streaming 10.21%11.09%16.49%3.06%
9 Zoom Scribe (Live) 10.50%15.95%8.09%7.47%
10 Mistral Voxtral Mini Realtime 12.26%12.67%16.02%8.09%
11 Deepgram Nova-3 13.22%19.01%17.37%3.29%
12 Speechmatics Standard Realtime 14.24%18.68%18.72%5.32%
13 Deepgram Flux EN 14.28%20.49%18.32%4.02%
14 Reson8 Reson8 14.31%16.37%21.81%4.76%
15 Azure Realtime STT 14.41%17.21%19.11%6.92%
16 xAI Grok Streaming 15.69%17.95%19.35%9.78%
17 Gladia Solaria 17.03%18.83%25.30%6.97%
18 OpenAI GPT Realtime 17.21%13.73%5.39%32.50%
19 Cartesia Ink-Whisper 21.09%24.39%31.09%7.78%
20 Smallest Pulse 21.40%25.42%30.77%8.01%
21 OpenAI GPT-4o Transcribe 26.56%23.34%14.91%41.43%
22 OpenAI GPT-4o mini Transcribe 27.94%25.45%19.27%39.11%

Best Others The best value in each column is filled blue.

Medication and disease

Table 5 Medication and disease terms scored separately. medication / disease entity error rate · runs 2026-08-18 → 2026-09-22 · one row per model, most recent run
Medication and disease terms scored separately.
# System model / snapshot Medication mean Primary care Dictation Respiratory Disease mean Primary care Dictation Respiratory
1 Speechmatics Enhanced Realtime 6.62%8.92%5.04%5.90%3.38%7.35%0.93%1.86%
2 Meta Muse Voice Transcribe 6.66%10.40%4.20%5.38%2.04%3.99%0.93%1.19%
3 AssemblyAI Universal-3.6 Pro Realtime 7.11%9.13%5.52%6.67%2.11%3.36%1.64%1.34%
4 Soniox Realtime 7.20%10.83%3.84%6.92%3.22%3.57%4.44%1.64%
5 AssemblyAI Universal-3.5 Pro Realtime 7.37%9.34%6.60%6.15%2.78%4.83%2.10%1.42%
6 ElevenLabs Scribe v2 Realtime 10.11%13.57%11.64%5.13%3.53%8.17%0.93%1.49%
7 Google Gemini 3.5 Transcribe Live 12.23%19.19%10.80%6.70%5.16%11.76%1.64%2.09%
8 AWS Transcribe Streaming 14.40%15.07%21.97%6.15%4.90%7.14%5.84%1.72%
9 Zoom Scribe (Live) 15.21%19.11%11.40%15.13%6.31%12.82%1.64%4.47%
10 Deepgram Nova-3 17.76%25.05%20.53%7.69%8.83%13.03%11.21%2.24%
11 OpenAI GPT Realtime 17.76%17.83%6.72%28.72%14.02%9.66%2.80%29.60%
12 Mistral Voxtral Mini Realtime 18.22%18.47%22.09%14.10%5.97%6.93%4.21%6.79%
13 Azure Realtime STT 18.60%23.57%21.97%10.26%9.57%10.92%13.55%4.25%
14 xAI Grok Streaming 19.50%24.63%26.05%7.81%7.74%11.34%6.31%5.56%
15 Deepgram Flux EN 19.77%29.30%20.77%9.23%9.18%11.76%13.55%2.24%
16 Speechmatics Standard Realtime 19.81%26.98%21.13%11.31%9.56%10.53%14.02%4.13%
17 Reson8 Reson8 21.23%23.78%27.85%12.05%7.23%9.00%10.05%2.61%
18 Gladia Solaria 23.23%23.81%32.41%13.46%9.69%12.86%11.45%4.75%
19 Smallest Pulse 26.83%31.55%41.42%7.53%12.04%18.30%10.05%7.76%
20 Cartesia Ink-Whisper 28.31%33.55%37.01%14.36%13.72%15.34%19.57%6.26%
21 OpenAI GPT-4o Transcribe 28.38%27.18%16.93%41.03%22.68%19.54%10.98%37.51%
22 OpenAI GPT-4o mini Transcribe 31.72%32.91%21.49%40.77%22.84%18.07%14.95%35.50%

Best Others The best value in each column is filled blue.

Data

○Primary-care consultations (PriMock57) 114 consultations
Public corpus of mock primary-care consultations: clinicians and actor-patients work through realistic appointments, recorded as telemedicine calls.
◆Medical dictation 93 recordings
Licensed clinician dictation, notes and findings spoken for the record, dense with medication and disease terms.
●Respiratory clinic recordings 272 recordings
Real clinical conversation from respiratory clinic encounters, with the accents, interruptions and room noise of a working clinic.

public, open source public, licensed private, ours

04English

General English: entity accuracy and non-speech handling first, then word error rate on short-form, accented and long-form sets.

Table 6 Judge-scored entity error (%) on general English, by class. We lead on both overall scores. Per class, we lead on person names, organisations and URLs. A 0.00 cell means no data. LLM-judge-scored entity error rate (normalized and formatted) · runs 2026-08-18 → 2026-09-22 · one row per model, most recent run
Judge-scored entity error (%) on general English, by class. We lead on both overall scores. Per class, we lead on person names, organisations and URLs. A 0.00 cell means no data.
System All classes All, formatted Persons Products Addresses Orgs Places Phone Email URLs Dates Credit cards Alphanum.
AssemblyAI Universal-3.5 Pro 18.127.419.919.221.516.420.59.316.715.116.519.69.0
AssemblyAI Universal-3.6 Pro 18.927.620.322.624.317.121.212.716.714.920.711.37.7
Google Gemini 3.5 Live 19.729.522.922.130.617.618.18.919.015.415.622.012.3
Speechmatics Enhanced 19.837.525.222.224.621.716.50.425.821.515.8n/d7.7
OpenAI GPT Realtime 20.238.624.827.521.420.017.72.119.921.919.03.013.4
ElevenLabs Scribe v2 20.429.224.822.629.222.618.9n/d17.214.919.4n/d8.7
Meta Muse Voice Transcribe 20.544.424.420.738.218.317.52.515.829.412.72.410.3
Mistral Voxtral Mini 20.831.727.824.026.419.317.43.813.130.016.4n/d7.5
Soniox Realtime 21.533.631.631.920.121.917.81.711.319.715.7—8.8
OpenAI GPT-4o Transcribe 24.437.826.225.933.023.326.218.125.319.721.921.410.5
Reson8 24.437.137.324.225.923.618.58.016.326.919.24.28.3
Gladia Solaria 24.941.232.526.629.224.122.14.221.331.020.27.710.3
Zoom Scribe (Live) 24.939.334.329.625.923.925.1—15.425.523.6—5.7
OpenAI GPT-4o mini Transcribe 25.038.128.727.133.124.027.28.924.020.120.28.312.6
Deepgram Nova-3 25.942.436.128.024.825.622.316.916.320.829.37.18.8
xAI Grok Streaming 26.150.234.028.827.624.826.12.519.928.324.24.88.5
Speechmatics Standard 26.644.433.925.529.126.220.46.330.839.816.5n/d26.5
Cartesia Ink-Whisper 27.843.037.628.232.927.326.112.718.924.220.913.710.7
AWS Transcribe 28.143.846.229.620.427.623.40.420.828.317.38.37.5
Deepgram Flux EN 31.462.549.133.032.730.225.03.415.831.423.5n/d10.0
Azure 33.197.545.035.950.234.427.82.114.524.731.00.610.0
Smallest Pulse 33.258.053.033.033.031.728.73.817.625.321.15.410.0

Best Worst Cells are coloured by rank within each column (solid blue is best, white is worst).

0 2 4 Mistral Voxtral Mini: 0.03% Mistral Voxtral Mini 0.03% OpenAI GPT Realtime: 0.07% OpenAI GPT Realtime 0.07% Deepgram Nova-3: 0.17% Deepgram Nova-3 0.17% Google Gemini 3.5 Live: 0.17% Google Gemini 3.5 Live 0.17% ElevenLabs Scribe v2: 0.20% ElevenLabs Scribe v2 0.20% Speechmatics Standard: 0.20% Speechmatics Standard 0.20% Azure: 0.24% Azure 0.24% Speechmatics Enhanced: 0.27% Speechmatics Enhanced 0.27% AssemblyAI Universal-3.6 Pro: 0.27% AssemblyAI Universal-3.6 Pro 0.27% AssemblyAI Universal-3.5 Pro: 0.27% AssemblyAI Universal-3.5 Pro 0.27% Meta Muse Voice Transcribe: 0.27% Meta Muse Voice Transcribe 0.27% Zoom Scribe (Live): 0.31% Zoom Scribe (Live) 0.31% OpenAI GPT-4o mini Transcribe: 0.37% OpenAI GPT-4o mini Transcribe 0.37% Soniox Realtime: 0.58% Soniox Realtime 0.58% Reson8: 0.68% Reson8 0.68% xAI Grok Streaming: 1.22% xAI Grok Streaming 1.22% AWS Transcribe: 1.25% AWS Transcribe 1.25% Smallest Pulse: 1.42% Smallest Pulse 1.42% OpenAI GPT-4o Transcribe: 2.07% OpenAI GPT-4o Transcribe 2.07% Cartesia Ink-Whisper: 3.15% Cartesia Ink-Whisper 3.15% Gladia Solaria: 4.88% Gladia Solaria 4.88% Deepgram Flux EN: 38.75% Deepgram Flux EN › 38.75%
Figure 6 Non-speech response rate: how often a system produces text for a clip with no speech. Lower is better. The axis is clipped at 5%; the notched bar exceeds it. non-speech response rate (share of pure non-speech clips that produce any text) · runs 2026-08-18 → 2026-09-22 · one row per model, most recent run

Word error rate — short-form

Table 7 Short-form English WER. We lead on the mean and Common Voice. Systems without a Common Voice run have no mean. normalized word error rate · runs 2026-08-18 → 2026-09-22 · one row per model, most recent run
Short-form English WER. We lead on the mean and Common Voice. Systems without a Common Voice run have no mean.
# System model / snapshot Mean 3 sets Common Voice LibriSpeech clean LibriSpeech other
1 AssemblyAI Universal-3.6 Pro Realtime 3.44%5.18%1.88%3.25%
2 Reson8 Reson8 3.70%6.98%1.43%2.68%
3 Zoom Scribe (Live) 3.75%6.32%1.72%3.21%
4 AssemblyAI Universal-3.5 Pro Realtime 3.83%6.44%1.83%3.22%
5 AWS Transcribe Streaming 3.93%6.12%1.90%3.77%
6 Speechmatics Enhanced Realtime 4.69%6.78%2.31%4.99%
7 Google Gemini 3.5 Transcribe Live 5.04%8.22%2.07%4.83%
8 ElevenLabs Scribe v2 Realtime 5.13%9.01%1.96%4.43%
9 Smallest Pulse 5.41%9.34%2.19%4.69%
10 Azure Realtime STT 5.54%8.68%2.44%5.49%
11 xAI Grok Streaming 5.81%11.04%2.04%4.36%
12 Meta Muse Voice Transcribe 6.20%11.24%2.15%5.22%
13 Speechmatics Standard Realtime 6.41%9.43%3.18%6.61%
14 Deepgram Flux EN 6.53%7.68%3.56%8.35%
15 Mistral Voxtral Mini Realtime 6.61%12.23%2.10%5.49%
16 Gladia Solaria 7.01%12.88%2.73%5.42%
17 Deepgram Nova-3 7.46%12.38%3.28%6.72%
18 Soniox Realtime 7.60%12.38%3.27%7.15%
19 Cartesia Ink-Whisper 7.66%13.31%3.19%6.48%
20 OpenAI GPT Realtime 7.95%13.83%2.59%7.44%
21 OpenAI GPT-4o Transcribe 8.11%14.99%2.20%7.15%
22 OpenAI GPT-4o mini Transcribe 8.75%15.76%2.53%7.97%

Best Others The best value in each column is filled blue.

Listen to the test data

Two read sentences from the short-form sets above, human reference beside each transcript. Words that differ from it are marked.

1 of 2
Short-form English17 s · Common Voice (CC0)

A read sentence. The other system appends a sentence and a name that are not in the audio.

Reference

You can do a lot just using your voice, but there are still a few times you'll find yourself reaching for a mouse.

AssemblyAI

You can do a lot just using your voice, but there are still a few times you'll find yourself reaching for a mouse.

Speechmatics Enhanced

I know you can do a lot just using your voice, but there are still a few times you'll find yourself reaching for a mouse. And I'm Larry Hanson. We're here in this beautiful hall.

Word error rate — accented English

Table 8 Accented English WER. We lead on British-accented English and Indian-accented English. normalized word error rate · runs 2026-08-18 → 2026-09-22 · one row per model, most recent run
Accented English WER. We lead on British-accented English and Indian-accented English.
# System model / snapshot Mean 4 sets British French-acc. Spanish-acc. Indian-acc.
1 Azure Realtime STT 6.09%6.25%4.44%5.52%8.15%
2 AssemblyAI Universal-3.6 Pro Realtime 7.29%5.53%8.54%9.62%5.47%
3 AssemblyAI Universal-3.5 Pro Realtime 7.46%5.72%8.69%9.53%5.89%
4 Zoom Scribe (Live) 7.94%6.14%9.57%10.03%6.04%
5 Reson8 Reson8 9.24%5.76%11.89%12.60%6.73%
6 Google Gemini 3.5 Transcribe Live 9.39%6.30%11.86%12.82%6.59%
7 Meta Muse Voice Transcribe 9.64%7.16%11.26%12.22%7.93%
8 xAI Grok Streaming 9.85%7.33%12.21%12.88%6.99%
9 ElevenLabs Scribe v2 Realtime 10.10%6.92%12.12%13.29%8.07%
10 OpenAI GPT-4o Transcribe 10.30%7.27%11.83%13.89%8.22%
11 Speechmatics Enhanced Realtime 10.44%6.70%13.26%14.80%6.98%
12 Gladia Solaria 11.07%8.67%12.60%13.85%9.15%
13 AWS Transcribe Streaming 11.10%7.02%15.49%15.47%6.42%
14 Cartesia Ink-Whisper 11.23%8.16%12.92%14.57%9.27%
15 OpenAI GPT-4o mini Transcribe 11.36%7.88%13.31%15.35%8.88%
16 Soniox Realtime 11.69%8.99%12.87%14.32%10.58%
17 OpenAI GPT Realtime 11.78%8.31%14.83%15.78%8.20%
18 Smallest Pulse 11.85%7.70%15.15%16.53%8.03%
19 Deepgram Nova-3 11.95%9.14%13.92%15.19%9.57%
20 Deepgram Flux EN 12.55%9.61%14.83%16.12%9.64%
21 Speechmatics Standard Realtime 12.58%8.43%16.29%17.46%8.14%
22 Mistral Voxtral Mini Realtime 12.63%7.39%14.19%15.87%13.06%

Best Others The best value in each column is filled blue.

Word error rate — long-form

Table 9 Long-form English WER. TED-LIUM references omit sponsor read-outs present in the audio. normalized word error rate · runs 2026-08-18 → 2026-09-22 · one row per model, most recent run
Long-form English WER. TED-LIUM references omit sponsor read-outs present in the audio.
# System model / snapshot Mean 3 sets TED-LIUM Rev16 CALLHOME per channel
1 Meta Muse Voice Transcribe 7.37%6.01%6.57%8.36%
2 Reson8 Reson8 7.74%6.01%6.87%9.17%
3 Speechmatics Enhanced Realtime 8.08%6.29%7.34%9.59%
4 Soniox Realtime 8.19%7.13%8.17%9.46%
5 AWS Transcribe Streaming 8.71%6.45%7.81%9.82%
6 AssemblyAI Universal-3.6 Pro Realtime 8.82%8.14%7.97%11.07%
7 Speechmatics Standard Realtime 8.90%6.83%8.06%10.77%
8 Azure Realtime STT 9.20%7.16%8.17%10.87%
9 Zoom Scribe (Live) 9.34%7.41%7.17%13.26%
10 AssemblyAI Universal-3.5 Pro Realtime 9.35%8.58%8.24%11.69%
11 Google Gemini 3.5 Transcribe Live 9.42%6.87%8.04%10.99%
12 xAI Grok Streaming 9.43%6.30%8.00%12.82%
13 Deepgram Flux EN 9.59%7.01%8.08%12.12%
14 Deepgram Nova-3 9.75%7.43%9.40%11.19%
15 Cartesia Ink-Whisper 10.60%7.13%10.42%15.62%
16 Gladia Solaria 10.72%7.75%9.04%18.25%
17 Smallest Pulse 10.95%5.23%7.62%14.13%
18 OpenAI GPT Realtime 15.87%5.70%6.02%20.90%
19 ElevenLabs Scribe v2 Realtime 17.12%5.96%8.38%45.51%
20 Mistral Voxtral Mini Realtime 18.93%6.96%8.33%52.91%
21 OpenAI GPT-4o mini Transcribe 19.10%7.28%21.26%18.56%
22 OpenAI GPT-4o Transcribe 27.45%9.88%28.51%30.33%

Best Others The best value in each column is filled blue.

Data

●Judge-scored entities
The short-form English audio below. An LLM judge grades each entity (person, product, organisation, place, address, phone, email, date, ID, URL) on whether the transcript names the same real-world thing, with and without formatting.
●Non-speech 2,950 clips
Hold music, line noise and silence, with no words at all. The score is how often a model emits text anyway.
○Common Voice 16,029 clips
Crowdsourced read sentences from Mozilla’s public Common Voice, recorded by volunteers on their own phones and laptops, so accents, microphones and rooms vary widely.
○LibriSpeech test-clean 2,620 utterances
Read English audiobook speech from the public LibriSpeech corpus, recorded close-mic by LibriVox volunteers. The "clean" split is the easier half: clear speakers, little noise.
○LibriSpeech test-other 2,939 utterances
The harder half of the same LibriSpeech set: speakers and recordings that automatic transcription finds more difficult, still read prose.
◆Accented English 9,315 files
Four licensed sets of prompted British-, French-, Spanish- and Indian-accented English: accuracy outside the General American accent most models are tuned on.
○TED-LIUM 11 full-length talks
Full TED talks from the public TED-LIUM corpus: one prepared speaker on stage, recorded end to end, many minutes of continuous speech per file.
○Rev16 9 full-length recordings
Podcast-style recordings from the public Rev16 set: spontaneous multi-speaker studio conversation, episode length.
◆CALLHOME 71 scored channels
Public CALLHOME telephone calls between friends and family: unscripted, fast, overlapping conversation over a real phone line. Each speaker's channel is scored on its own.

public, open source public, licensed private, ours

05Multilingual

Eighteen languages on FLEURS and Common Voice, drawn from the thirty-two Universal-3.6 Pro Realtime supports. A system is scored on every language it supports, and a cross marks one its vendor does not document for realtime use — so the means below cover different numbers of languages and the Languages column says how many.

FLEURS — Systems covering all 18 languages

Table 10 WER on FLEURS (CER for languages marked *). Each mean covers only the languages that system supports; the Languages column gives the count. normalized word error rate (character error rate for Japanese, Chinese, Korean) · runs 2026-08-25 → 2026-09-23 · one row per model, most recent run
WER on FLEURS (CER for languages marked *). Each mean covers only the languages that system supports; the Languages column gives the count.
System Mean over supported langs Languages of 18 Spanish French German Italian Portuguese Dutch Danish Swedish Finnish Russian Turkish Arabic Hebrew Hindi Vietnamese Japanese* Chinese* Korean*
Google Gemini 3.5 Live 6.4183.85.44.73.24.75.09.17.37.55.35.36.718.95.05.25.09.43.8
Soniox Realtime 7.9185.07.86.94.96.27.711.410.38.07.37.37.219.26.37.05.810.34.4
AssemblyAI Universal-3.6 Pro 8.4184.15.05.04.05.55.912.76.610.87.58.612.118.27.513.96.314.04.3
AWS Transcribe 8.9185.37.68.04.68.27.111.510.09.89.48.68.522.66.39.37.012.34.9
ElevenLabs Scribe v2 9.3184.66.95.84.36.47.012.510.99.56.67.710.827.015.88.88.610.14.6
Meta Muse Voice Transcribe 10.0186.58.06.46.08.78.613.611.010.710.110.58.927.08.16.59.813.86.1
Speechmatics Enhanced 11.1184.76.56.85.47.38.511.89.98.010.09.16.518.57.08.631.333.85.4
Speechmatics Standard 11.6185.48.18.25.77.39.511.710.28.710.58.78.419.37.09.331.434.75.2
OpenAI GPT Realtime 12.3185.711.811.25.98.810.114.712.412.410.414.310.534.510.411.413.416.37.1
Azure 12.7187.711.49.18.310.912.919.816.515.613.115.212.527.68.69.48.611.59.4
OpenAI GPT-4o Transcribe 12.8184.310.613.15.58.513.017.710.19.28.414.89.434.712.418.716.117.46.7
OpenAI GPT-4o mini Transcribe 15.6185.513.214.96.910.916.820.612.712.49.917.511.745.016.823.115.419.37.6
Gladia Solaria 16.2186.712.710.97.58.711.622.615.315.910.312.215.145.423.921.523.621.16.6
Cartesia Ink-Whisper 18.7189.016.912.79.710.717.425.416.117.015.716.218.142.027.523.620.426.310.9

Best Worst Cells are coloured by rank within each column (solid blue is best, white is worst).

FLEURS — Systems covering fewer than 18

Table 11 The same measurement on the systems that do not cover all eighteen. Each mean is over that system’s own languages, so these are not comparable with the table above or with each other — the Languages column gives the count, and a cross marks a language the vendor does not document for realtime use. normalized word error rate (character error rate for Japanese, Chinese, Korean) · runs 2026-08-25 → 2026-09-23 · one row per model, most recent run
The same measurement on the systems that do not cover all eighteen. Each mean is over that system’s own languages, so these are not comparable with the table above or with each other — the Languages column gives the count, and a cross marks a language the vendor does not document for realtime use.
System Mean over supported langs Languages of 18 Spanish French German Italian Portuguese Dutch Danish Swedish Finnish Russian Turkish Arabic Hebrew Hindi Vietnamese Japanese* Chinese* Korean*
Reson8 5.874.05.55.34.05.66.3✕9.6✕✕✕✕✕✕✕✕✕✕
Zoom Scribe (Live) 6.275.25.96.55.26.7✕✕✕✕✕✕7.9✕✕✕6.2—✕
Mistral Voxtral Mini 9.5124.89.67.85.56.99.5———8.2—12.7—16.8—9.616.55.8
AssemblyAI Universal-3.5 Pro 9.9164.45.25.34.35.86.513.17.512.5—9.814.021.912.415.96.513.4—
Smallest Pulse 11.9118.912.711.67.011.313.1—✕✕15.8———10.7—12.720.17.4
Deepgram Flux 12.797.813.910.28.812.215.8✕✕✕11.5✕✕✕20.1✕13.6✕✕
Deepgram Nova-3 14.298.914.013.310.013.314.6———12.5———22.3—18.8——
xAI Grok Streaming 14.8146.126.010.96.28.412.528.323.2✕12.015.415.4✕—12.921.7✕7.8

Best Worst Cells are coloured by rank within each column (solid blue is best, white is worst).

Common Voice — Systems covering all 18 languages

Table 12 WER on Common Voice (CER for languages marked *). Each mean covers only the languages that system supports; the Languages column gives the count. normalized word error rate (character error rate for Japanese, Chinese, Korean) · runs 2026-08-25 → 2026-09-23 · one row per model, most recent run
WER on Common Voice (CER for languages marked *). Each mean covers only the languages that system supports; the Languages column gives the count.
System Mean over supported langs Languages of 18 Spanish French German Italian Portuguese Dutch Danish Swedish Finnish Russian Turkish Arabic Hebrew Hindi Vietnamese Japanese* Chinese* Korean*
Speechmatics Enhanced 7.3182.57.23.74.75.22.47.34.03.47.17.36.19.43.92.823.225.17.2
AWS Transcribe 7.5184.18.85.84.24.92.77.35.95.74.810.39.615.05.18.215.98.57.5
AssemblyAI Universal-3.6 Pro 7.6183.26.93.63.34.92.69.44.89.38.111.09.412.35.310.516.29.26.8
Speechmatics Standard 9.0184.210.26.25.36.43.08.15.75.97.76.511.111.84.06.225.425.88.3
Google Gemini 3.5 Live 9.2184.09.35.24.25.83.311.59.711.56.510.19.617.84.79.321.713.37.5
Soniox Realtime 9.5186.012.27.66.37.44.69.58.67.79.010.810.212.56.610.119.214.18.7
Azure 10.0186.811.06.16.27.24.611.910.011.17.012.311.520.04.98.819.89.910.4
Meta Muse Voice Transcribe 10.2184.910.46.65.56.45.714.611.113.67.29.69.318.45.89.020.316.49.6
OpenAI GPT Realtime 12.3186.912.69.37.89.46.814.512.011.88.211.112.728.98.413.420.216.410.5
ElevenLabs Scribe v2 14.6185.810.27.16.19.75.514.512.914.610.215.020.643.421.621.520.712.111.0
OpenAI GPT-4o Transcribe 15.0188.915.510.39.213.27.818.012.611.29.315.117.234.910.020.822.021.412.1
Gladia Solaria 16.1187.814.410.59.89.55.717.512.614.09.516.219.836.122.020.626.229.09.6
OpenAI GPT-4o mini Transcribe 17.7189.816.811.710.614.311.921.917.113.311.914.322.542.921.523.922.319.312.2
Cartesia Ink-Whisper 18.91810.417.712.411.111.59.721.116.317.712.318.022.039.127.922.228.329.213.8

Best Worst Cells are coloured by rank within each column (solid blue is best, white is worst).

Common Voice — Systems covering fewer than 18

Table 13 The same measurement on the systems that do not cover all eighteen. Each mean is over that system’s own languages, so these are not comparable with the table above or with each other — the Languages column gives the count, and a cross marks a language the vendor does not document for realtime use. normalized word error rate (character error rate for Japanese, Chinese, Korean) · runs 2026-08-25 → 2026-09-23 · one row per model, most recent run
The same measurement on the systems that do not cover all eighteen. Each mean is over that system’s own languages, so these are not comparable with the table above or with each other — the Languages column gives the count, and a cross marks a language the vendor does not document for realtime use.
System Mean over supported langs Languages of 18 Spanish French German Italian Portuguese Dutch Danish Swedish Finnish Russian Turkish Arabic Hebrew Hindi Vietnamese Japanese* Chinese* Korean*
Reson8 4.973.57.24.13.14.32.9✕9.2✕✕✕✕✕✕✕✕✕✕
Zoom Scribe (Live) 8.075.78.65.84.16.6✕✕✕✕✕✕8.6✕✕✕16.4—✕
AssemblyAI Universal-3.5 Pro 10.5164.38.35.04.48.03.611.06.815.2—15.514.114.910.613.819.613.1—
Smallest Pulse 12.4118.612.311.36.621.69.2—✕✕10.8———9.7—17.518.010.7
Deepgram Flux 13.697.813.79.910.810.59.1✕✕✕13.1✕✕✕22.7✕24.4✕✕
Mistral Voxtral Mini 13.7125.711.58.87.49.98.5———9.8—17.4—18.1—22.830.514.3
Deepgram Nova-3 15.298.413.812.810.812.79.3———14.4———25.2—29.7——
xAI Grok Streaming 17.3148.913.711.79.111.19.530.828.8✕13.723.724.2✕—17.825.0✕13.5

Best Worst Cells are coloured by rank within each column (solid blue is best, white is worst).

Data

○FLEURS
Public corpus of parallel sentences read by native speakers in quiet settings, one test set per language. Shown: Spanish, French, German, Italian, Portuguese, Dutch, Hindi, Japanese and Chinese.
○Common Voice
Mozilla’s public Common Voice, per language: volunteers reading sentences on their own devices, so accents and microphones vary. Japanese and Chinese are scored by character error rate.

public, open source public, licensed private, ours

06Code-switching

Speech that switches between English and a second language within one utterance: five language pairs plus a Spanish–English conversational corpus.

Table 14 Code-switching WER. We lead on the five-pair mean. We lead in all 5 pairs. normalized word error rate · runs 2026-08-18 → 2026-09-23 · one row per model, most recent run · pair value = mean of 4 sub-sets
Code-switching WER. We lead on the five-pair mean. We lead in all 5 pairs.
# System model / snapshot Mean 5 pairs EN↔ES EN↔DE EN↔FR EN↔IT EN↔PT Miami 10h
1 AssemblyAI Universal-3.6 Pro Realtime 7.20%9.59%6.13%7.23%6.03%7.00%24.28%
2 AssemblyAI Universal-3.5 Pro Realtime 8.65%10.95%7.84%8.54%7.21%8.71%27.31%
3 AWS Transcribe Streaming 9.13%11.32%8.12%9.33%7.05%9.81%27.96%
4 Google Gemini 3.5 Transcribe Live 10.16%14.40%8.29%10.05%8.25%9.79%40.69%
5 Reson8 Reson8 10.67%12.58%9.73%12.60%8.77%9.67%24.48%
6 Soniox Realtime 11.26%12.12%10.81%12.15%10.04%11.18%21.93%
7 ElevenLabs Scribe v2 Realtime 13.01%14.27%12.29%13.16%11.53%13.79%27.09%
8 Meta Muse Voice Transcribe 13.34%13.34%12.31%14.64%11.56%14.85%23.10%
9 Smallest Pulse 14.56%18.52%13.14%13.61%10.79%16.73%47.17%
10 Azure Realtime STT 14.60%17.25%13.33%14.86%12.52%15.03%34.17%
11 Gladia Solaria 17.09%18.39%18.10%18.73%14.28%15.93%34.49%
12 Deepgram Nova-3 Multi 17.13%16.66%18.01%18.52%15.17%17.29%29.91%
13 Mistral Voxtral Mini Realtime 18.46%18.37%19.11%20.79%16.61%17.43%34.35%
14 xAI Grok Streaming 20.61%21.73%20.06%21.69%18.79%20.79%34.20%
15 Zoom Scribe (Live) 32.48%26.66%35.09%32.56%32.09%36.02%37.87%
16 OpenAI GPT-4o mini Transcribe 33.27%35.28%33.06%34.22%31.73%32.07%45.01%
17 OpenAI GPT-4o Transcribe 35.46%39.66%34.70%35.78%33.17%34.00%61.27%
18 OpenAI GPT Realtime 37.35%39.34%35.46%37.87%36.30%37.77%53.31%
19 Speechmatics Enhanced Realtime 38.23%33.75%33.78%35.98%45.39%42.23%36.04%
20 Speechmatics Standard Realtime 42.86%37.60%44.12%42.99%45.11%44.49%40.95%
21 Cartesia Ink-Whisper EN 50.52%46.90%51.50%53.10%53.17%47.95%43.72%
22 Deepgram Flux EN 52.86%52.36%52.33%56.46%53.39%49.76%45.55%
23 Deepgram Nova-3 EN 54.71%54.39%54.61%58.15%56.24%50.15%43.85%

Best Others The best value in each column is filled blue.

0 20 40 AssemblyAI Universal-3.6 Pro: 7.20% AssemblyAI Universal-3.6 Pro 7.20% AssemblyAI Universal-3.5 Pro: 8.65% AssemblyAI Universal-3.5 Pro 8.65% AWS Transcribe: 9.13% AWS Transcribe 9.13% Google Gemini 3.5 Live: 10.16% Google Gemini 3.5 Live 10.16% Reson8: 10.67% Reson8 10.67% Soniox Realtime: 11.26% Soniox Realtime 11.26% ElevenLabs Scribe v2: 13.01% ElevenLabs Scribe v2 13.01% Meta Muse Voice Transcribe: 13.34% Meta Muse Voice Transcribe 13.34% Smallest Pulse: 14.56% Smallest Pulse 14.56% Azure: 14.60% Azure 14.60% Gladia Solaria: 17.09% Gladia Solaria 17.09% Deepgram Nova-3 Multi: 17.13% Deepgram Nova-3 Multi 17.13% Mistral Voxtral Mini: 18.46% Mistral Voxtral Mini 18.46% xAI Grok Streaming: 20.61% xAI Grok Streaming 20.61% Zoom Scribe (Live): 32.48% Zoom Scribe (Live) 32.48% OpenAI GPT-4o mini Transcribe: 33.27% OpenAI GPT-4o mini Transcribe 33.27% OpenAI GPT-4o Transcribe: 35.46% OpenAI GPT-4o Transcribe 35.46% OpenAI GPT Realtime: 37.35% OpenAI GPT Realtime 37.35% Speechmatics Enhanced: 38.23% Speechmatics Enhanced 38.23% Speechmatics Standard: 42.86% Speechmatics Standard 42.86% Cartesia Ink-Whisper EN: 50.52% Cartesia Ink-Whisper EN 50.52% Deepgram Flux EN: 52.86% Deepgram Flux EN 52.86% Deepgram Nova-3 EN: 54.71% Deepgram Nova-3 EN 54.71%
Figure 7 Five-pair mean WER. normalized word error rate · runs 2026-08-18 → 2026-09-23 · one row per model, most recent run

Data

○English ↔ Spanish, German, French, Italian, Portuguese 5 pairs · 4 sets each
Audio that alternates between English and the second language mid-stream, built from public read-speech corpora in short, long and split-long clips so the switch falls at different points.
◆Miami, 10-hour selection 21 recordings
A ten-hour selection of the Miami Spanish-English conversation, with references produced independently by a third-party transcription service.

public, open source public, licensed private, ours

07Diarization

Speaker attribution on CALLHOME, AMI, DiPCo and NOTSOFAR, for systems that offer streaming diarization. Diarization error rate is the share of audio time given to the wrong speaker, missed, or wrongly marked as speech.

Table 15 Diarization error rate and cpWER. Reson8 caps speaker count at 4, its product maximum. That fits CALLHOME, AMI and DiPCo, whose references never exceed four speakers, and costs it on NOTSOFAR, which runs to seven, so its per-corpus figures are not directly comparable with systems that estimate speaker count freely. diarization error rate (DER) and cpWER · runs 2026-08-20 → 2026-09-24 · one row per model, most recent run
Diarization error rate and cpWER. Reson8 caps speaker count at 4, its product maximum. That fits CALLHOME, AMI and DiPCo, whose references never exceed four speakers, and costs it on NOTSOFAR, which runs to seven, so its per-corpus figures are not directly comparable with systems that estimate speaker count freely.
# System model / snapshot DER mean 4 sets CALLHOME AMI DiPCo NOTSOFAR cpWER mean CALLHOME AMI DiPCo NOTSOFAR
1 Meta Muse Voice Transcribe 14.51%16.44%19.45%14.81%7.34%24.23%12.54%22.66%26.65%35.07%
2 AssemblyAI Universal-3.5 Pro Realtime 15.96%13.87%18.31%21.29%10.37%30.59%19.76%28.00%34.00%40.61%
3 AssemblyAI Universal-3.6 Pro Realtime 16.68%13.74%18.79%22.10%12.09%30.92%19.21%28.16%34.20%42.11%
4 Reson8 Reson8 16.72%15.99%15.93%15.98%18.98%34.32%22.34%28.21%34.61%52.11%
5 Azure Realtime STT 20.43%20.21%18.00%24.30%19.19%40.96%34.68%33.63%41.45%54.06%
6 Speechmatics Enhanced Realtime 22.16%14.58%19.96%28.37%25.74%37.14%24.69%23.01%42.59%58.25%
7 xAI Grok Streaming 25.68%15.59%31.90%35.28%19.96%41.81%25.77%37.95%44.72%58.82%
8 Speechmatics Standard Realtime 25.71%25.43%21.27%24.68%31.47%48.86%47.84%33.66%43.43%70.49%
9 AWS Transcribe Streaming 27.48%27.89%26.16%32.09%23.79%46.07%41.66%38.72%47.91%55.98%
10 Soniox Realtime 31.77%19.83%46.27%27.23%33.76%42.78%17.20%55.47%39.56%58.88%
11 Deepgram Nova-3 34.01%30.46%33.86%45.35%26.37%56.09%52.17%43.99%69.88%58.30%
12 Smallest Pulse 41.35%27.16%59.41%50.98%27.87%53.34%31.65%65.89%62.56%53.26%

Best Others The best value in each column is filled blue.

Data

◆CALLHOME 40 calls
Public CALLHOME telephone calls between friends and family: unscripted, fast, heavily overlapping conversation over a real phone line.
○AMI 24 meetings
Public AMI Meeting Corpus: small groups of colleagues in real design and project meetings, in instrumented rooms.
○NOTSOFAR 129 meetings
Public NOTSOFAR corpus: real multi-party meetings on a single distant microphone in ordinary meeting rooms.
○DiPCo 10 sessions
Public Dinner Party Corpus: groups talking over dinner on far-field microphones, constant overlap and background clatter.

public, open source public, licensed private, ours

08Independent evaluations

Third-party results on their own corpora and pipelines, reproduced for comparison.

Artificial Analysis — Speech to Text (Streaming)

AA-WER v2.2 (May 2026): 50% AA-AgentTalk, 25% VoxPopuli, 25% Earnings22. End of speech is detected with SileroVAD; latency runs from there to the final transcript. Price is per 1,000 minutes of audio, normalized across billing models. The leaderboard publishes no “last updated” date. The two AssemblyAI configurations shown are Universal-3.5 Pro Realtime. Universal-3.6 Pro has not been scored here yet; this section will be updated once the source publishes it.

Table 16 30 of the 31 systems on the Artificial Analysis streaming leaderboard, at published ranks; our earlier-generation model is omitted. source = https://artificialanalysis.ai/speech-to-text/streaming · read 2026-09-09 · price shown for every system the source publishes one for
30 of the 31 systems on the Artificial Analysis streaming leaderboard, at published ranks; our earlier-generation model is omitted.
# System AA-WER index % t final s Price $/1k min
1 Muse Voice Transcribe (Meta) 3.06%0.163 s3.0
2 Cartesia Ink-2 (semantic) 3.36%0.431 s4.0
3 ElevenLabs Scribe v2 Realtime 3.59%0.141 s6.5
4 Qwen3 ASR Flash Realtime 3.73%0.476 s5.4
5 GPT Live Transcribe 3.92%0.812 s17.0
6 Grok STT Streaming 3.93%0.373 s3.3
7 Gemini 3.5 Transcribe Live 4.00%0.395 s9.0
8 AssemblyAI U3.5 Realtime Pro — min latency 4.02%0.191 s7.5
9 Cartesia Ink-2 (external endpoint) 4.02%0.067 s4.0
10 AssemblyAI U3.5 Realtime Pro — max accuracy 4.03%0.202 s7.5
12 Inworld STT 1 4.18%0.071 s1.4
13 Soniox v4 4.49%0.062 s2.0
14 Soniox v5 4.50%0.054 s2.0
15 Google Chirp 3 Streaming 4.80%1.276 s16.0
16 OpenAI GPT Realtime (Whisper) 4.89%0.688 s17.0
17 Smallest Pulse 4.98%0.132 s4.0
18 Voxtral Mini Realtime 5.24%0.682 s6.0
19 Azure Speech 5.25%0.625 s16.7
20 Nemotron 3 ASR (1120 ms) 5.36%0.418 s—
21 Amazon Transcribe 5.79%0.620 s24.0
22 Nemotron 3 ASR (560 ms) 5.96%0.252 s—
23 Deepgram Nova-3 Realtime 6.59%0.066 s4.8
24 Gradium 6.61%0.435 s12.0
25 Nemotron 3 ASR (80 ms, Together) 7.03%0.070 s—
26 Deepgram Flux 7.39%0.021 s6.5
27 Nemotron 3 ASR (160 ms) 7.61%0.098 s—
28 Gladia Solaria 7.82%1.523 s12.5
29 Speechmatics Realtime Enhanced 8.05%0.317 s17.5
30 Nemotron 3 ASR (80 ms) 8.38%0.073 s—
31 Rev.ai 10.14%0.697 s3.3

Best Others The best value in each column is filled blue.

Table 17 The three sets in the Artificial Analysis index. The source ranks our two configurations separately; both are shown beside each set’s leading system. source = https://artificialanalysis.ai/speech-to-text/streaming · read 2026-09-09
The three sets in the Artificial Analysis index. The source ranks our two configurations separately; both are shown beside each set’s leading system.
SetOur configurationWERRankLeading system on the setIts WER
VoxPopuliU3.5 RT Pro — min latency2.06%3 of 31Gemini 3.5 Transcribe Live1.76%
VoxPopuliU3.5 RT Pro — max accuracy2.09%4 of 31Gemini 3.5 Transcribe Live1.76%
AA-AgentTalkU3.5 RT Pro — min latency3.53%9 of 31Qwen3 ASR Flash Realtime2.49%
AA-AgentTalkU3.5 RT Pro — max accuracy3.48%7 of 31Qwen3 ASR Flash Realtime2.49%
Earnings22U3.5 RT Pro — min latency6.96%11 of 31Muse Voice Transcribe4.31%
Earnings22U3.5 RT Pro — max accuracy7.08%12 of 31Muse Voice Transcribe4.31%
0.02 0.05 0.1 0.2 0.5 1 2% 4% 6% 8% 10% Time to final transcript after end of speech (s) AA-WER index (%) Muse Voice Transcribe (Meta) — 0.163 s · 3.06% Muse Voice Transcribe (Meta) 0.163 s · 3.06% Cartesia Ink-2 (semantic) — 0.431 s · 3.36% Cartesia Ink-2 (semantic) 0.431 s · 3.36% ElevenLabs Scribe v2 Realtime — 0.141 s · 3.59% 1 Qwen3 ASR Flash Realtime — 0.476 s · 3.73% Qwen3 ASR Flash Realtime 0.476 s · 3.73% GPT Live Transcribe — 0.812 s · 3.92% GPT Live Transcribe 0.812 s · 3.92% Grok STT Streaming — 0.373 s · 3.93% Grok STT Streaming 0.373 s · 3.93% Gemini 3.5 Transcribe Live — 0.395 s · 4.00% 2 U3.5 RT Pro — min latency — 0.191 s · 4.02% U3.5 RT Pro — min latency 0.191 s · 4.02% Cartesia Ink-2 (external endpoint) — 0.067 s · 4.02% Cartesia Ink-2 (external endpoint) 0.067 s · 4.02% U3.5 RT Pro — max accuracy — 0.202 s · 4.03% U3.5 RT Pro — max accuracy 0.202 s · 4.03% Inworld STT 1 — 0.071 s · 4.18% Inworld STT 1 0.071 s · 4.18% Soniox v4 — 0.062 s · 4.49% Soniox v4 0.062 s · 4.49% Soniox v5 — 0.054 s · 4.50% Soniox v5 0.054 s · 4.50% Google Chirp 3 Streaming — 1.276 s · 4.80% 3 OpenAI GPT Realtime (Whisper) — 0.688 s · 4.89% 4 Smallest Pulse — 0.132 s · 4.98% Smallest Pulse 0.132 s · 4.98% Voxtral Mini Realtime — 0.682 s · 5.24% Voxtral Mini Realtime 0.682 s · 5.24% Azure Speech — 0.625 s · 5.25% Azure Speech 0.625 s · 5.25% Nemotron 3 ASR (1120 ms) — 0.418 s · 5.36% Nemotron 3 ASR (1120 ms) 0.418 s · 5.36% Amazon Transcribe — 0.62 s · 5.79% Amazon Transcribe 0.62 s · 5.79% Nemotron 3 ASR (560 ms) — 0.252 s · 5.96% 5 Deepgram Nova-3 Realtime — 0.066 s · 6.59% Deepgram Nova-3 Realtime 0.066 s · 6.59% Gradium — 0.435 s · 6.61% Gradium 0.435 s · 6.61% Nemotron 3 ASR (80 ms, Together) — 0.07 s · 7.03% Nemotron 3 ASR (80 ms, Together) 0.07 s · 7.03% Deepgram Flux — 0.021 s · 7.39% Deepgram Flux 0.021 s · 7.39% Nemotron 3 ASR (160 ms) — 0.098 s · 7.61% Nemotron 3 ASR (160 ms) 0.098 s · 7.61% Gladia Solaria — 1.523 s · 7.82% Gladia Solaria 1.523 s · 7.82% Speechmatics Realtime Enhanced — 0.317 s · 8.05% Speechmatics Realtime Enhanced 0.317 s · 8.05% Nemotron 3 ASR (80 ms) — 0.073 s · 8.38% Nemotron 3 ASR (80 ms) 0.073 s · 8.38% Rev.ai — 0.697 s · 10.14% Rev.ai 0.697 s · 10.14%
  1. 1 ElevenLabs Scribe v2 Realtime · 0.141 s · 3.59%
  2. 2 Gemini 3.5 Transcribe Live · 0.395 s · 4.00%
  3. 3 Google Chirp 3 Streaming · 1.276 s · 4.80%
  4. 4 OpenAI GPT Realtime (Whisper) · 0.688 s · 4.89%
  5. 5 Nemotron 3 ASR (560 ms) · 0.252 s · 5.96%
Figure 8 AA-WER against time to final transcript. The dashed line is the accuracy–latency frontier. source = https://artificialanalysis.ai/speech-to-text/streaming

Coval — Live STT Leaderboard

7-day window, read 2026-09-28 (snapshot generation 1045). The source’s default view is a rolling 24 hours; this is the wider window. WER is pooled over seven datasets — PipeCat production speech (897 clips) and six WildASR conditions (accents 845 clips, clean, clipping, far-field, noise gaps and phone codec at 284 each). Each paired WildASR condition reuses the clean baseline’s utterance, so a difference isolates the condition. Latency is the per-model average. The PipeCat references are model-generated rather than human-verified, which the source notes makes that set better for comparing models than for reading an absolute error rate. Universal-3.6 Pro’s three figures were supplied ahead of the source’s own publication of that model; its placings are computed by inserting them into the 2026-09-28 ranking rather than read from it, and should be re-read once the source publishes. Coval sells voice-agent evaluation tooling.

Table 18 Selected rows of 30. Shown as published, except Universal-3.6 Pro, whose figures precede the source’s publication of that model. Three AssemblyAI models are listed. source = https://benchmarks.coval.ai/stt · read 2026-09-28 · 7-day window
Selected rows of 30. Shown as published, except Universal-3.6 Pro, whose figures precede the source’s publication of that model. Three AssemblyAI models are listed.
System model / snapshot WER pooled %
AssemblyAI Universal 3.6 Pro rank 1 of 30 2.30%
AssemblyAI Universal 3.5 Pro rank 2 of 30 2.48%
Nari Labs Qwen3 ASR Fast rank 3 of 30 2.72%
Reson8 resonant-1 rank 4 of 30 2.76%
Baseten Qwen3 ASR 1.7b rank 5 of 30 2.78%
Gemini 3.5 Transcribe Live rank 6 of 30 3.01%
ElevenLabs Scribe v2 Realtime rank 17 of 30 4.47%
Soniox STT RT v5 rank 19 of 30 4.80%
Deepgram Flux rank 22 of 30 5.77%
AssemblyAI Universal Streaming rank 23 of 30 5.78%

Best Others The best value in each column is filled blue.

Figure 9 Universal-3.5 Pro’s WER by recording condition — the source publishes no per-condition breakdown for Universal-3.6 Pro. The dashed line is Universal-3.5 Pro’s own pooled value from Table 18. source = https://benchmarks.coval.ai/stt
0 1 2 3 4 PipeCat (production): 1.80% PipeCat (production) 1.80% WildASR accents: 2.31% WildASR accents 2.31% WildASR clean: 3.10% WildASR clean 3.10% WildASR phone codec: 3.53% WildASR phone codec 3.53% WildASR noise gaps: 3.72% WildASR noise gaps 3.72% WildASR clipping: 3.84% WildASR clipping 3.84% WildASR far-field: 4.11% WildASR far-field › 4.11% pooled 2.48%

Pipecat / Daily — STT benchmark

1,000 samples from the smart-turn-data-v3.1-train dataset; the source does not describe the set as English. Pooled WER is total errors over total reference words, judged semantically; TTFS is the median and P95 from when the speaker stops to the final transcription segment. Reproduced with the source’s own model names; ranked here by Pooled WER, where the source now lists rows alphabetically by vendor. Read from the repository README (last changed 2026-09-18). The source also publishes a per-sample WER mean, a transcript count, a perfect-transcript count and a P99, which are not carried here.

Table 19 Shown as published. Three AssemblyAI entries are listed; the source marks universal-3-6-pro as the current one and the other two as superseded. source = https://github.com/pipecat-ai/stt-benchmark · read 2026-09-28
Shown as published. Three AssemblyAI entries are listed; the source marks universal-3-6-pro as the current one and the other two as superseded.
# System Pooled WER % TTFS median ms TTFS p95 ms
1 Meta muse-voice-transcribe-1.0 0.83%3921292
2 AssemblyAI universal-3-6-pro 0.96%307401
3 Speechmatics linden-1 1.05%369438
4 Speechmatics 1.07%495676
5 Azure 1.18%10161345
6 AssemblyAI universal-3-5-pro 1.22%282354
7 Cartesia ink-2 1.25%299328
8 Soniox stt-rt-v5 1.27%260305
9 Soniox stt-rt-v4 1.29%249281
10 Deepgram nova-3-general 1.62%247298
11 AWS 1.75%11361527
12 NVIDIA Nemotron 3.0 ASR (en) 1.95%221238
13 Google gemini-3.5-transcribe-live 2.24%458532
14 Smallest AI pulse 2.37%398533
15 OpenAI gpt-realtime-whisper 2.73%740878
16 Google latest-long 2.85%8781155
17 AssemblyAI universal-streaming-english 3.02%256362
18 OpenAI gpt-4o-transcribe 3.06%637965
19 ElevenLabs scribe_v2_realtime 3.12%281348
20 Gradium default 3.71%570596
21 Cartesia ink-whisper 4.36%266364
22 NVIDIA Nemotron 3.5 ASR (multilingual) 4.58%236253
23 Mistral voxtral-mini-transcribe-realtime-2602 4.97%525973

Best Others The best value in each column is filled blue.

09How we measure

How the results were produced: the harness, the metrics, why public leaderboards can rank the same systems differently, and the limits of this report.

How we run every system

Protocol. Audio is streamed in real time to each system’s public API, default settings, language fixed. One normalizer is applied before scoring. Each streaming model a vendor ships has its own row.

Complete runs only. An average is shown only when the system was scored on every dataset in it; partial rows keep their cells but are not ranked.

n/d, n/s. A 0.00 error rate means no data. WER or CER of 60% or more means the language is unsupported; the cell is excluded from means.

Excluded. Unreleased models, and results on a superseded version of the voice-agent test set.

Systems

Every system with a public streaming API and a current run; a dot marks each suite it was scored on.

Table 20 Systems evaluated and the suites each has a current run on. 5 suites · runs 2026-08-27 → 2026-09-22 · added this update: Meta, Zoom
Systems evaluated and the suites each has a current run on.
VendorModelEnglish / CIVoice agentsCode-switch.DiarizationMultilingualRuns
AssemblyAIUniversal-3.5 Pro Realtime●●●●●2026-08-18 → 2026-09-02
AssemblyAIUniversal-3.6 Pro Realtime●●●●●2026-09-03 → 2026-09-21
AWSTranscribe Streaming●●●●●2026-08-19 → 2026-08-31
AzureAzure Realtime STT●●●●●2026-08-19 → 2026-09-11
CartesiaInk 2 (turns)·····2026-08-31
CartesiaInk-Whisper●●·✕●2026-08-19 → 2026-08-30
CartesiaInk-Whisper EN··●✕·2026-08-20
DeepgramFlux···✕●2026-08-25
DeepgramFlux EN●●●✕·2026-08-19 → 2026-09-11
DeepgramFlux Multi·●·✕·2026-08-27
DeepgramNova-3●●·●●2026-08-18 → 2026-08-31
DeepgramNova-3 EN··●··2026-08-20
DeepgramNova-3 Multi··●··2026-08-20
ElevenLabsScribe v2 Realtime●●●✕●2026-08-19 → 2026-08-31
GladiaSolaria●●●✕●2026-08-19 → 2026-09-10
GoogleGemini 3.5 Transcribe Live●●●✕●2026-08-31 → 2026-09-06
MetaMuse Voice Transcribe (new)●●●●●2026-09-03 → 2026-09-15
MistralVoxtral Mini Realtime●●●✕●2026-08-19 → 2026-08-31
OpenAIGPT Realtime●·●✕●2026-08-19 → 2026-08-31
OpenAIGPT-4o mini Transcribe●●●✕●2026-08-19 → 2026-08-31
OpenAIGPT-4o Transcribe●●●✕●2026-08-19 → 2026-08-31
Reson8Reson8●●●●●2026-09-01 → 2026-09-18
SmallestPulse●●●●●2026-08-28 → 2026-08-31
SonioxSoniox Realtime●●●●●2026-08-31 → 2026-09-23
SpeechmaticsEnhanced Realtime●●●●●2026-08-19 → 2026-08-31
SpeechmaticsStandard Realtime●●●●●2026-08-19 → 2026-08-31
xAIGrok Streaming●●●●●2026-08-19 → 2026-09-14
ZoomScribe (Live) (new)●●●✕●2026-09-10 → 2026-09-16

Test data

The audio each benchmark uses and the metric it scores. Each results section also lists its own datasets.

Table 21 Test data by benchmark. Files are scored recordings or clips; multilingual, code-switching and judge-scored sets are scored per language or label, so no single count applies. Public corpora are named, licensed sets described.
Test data by benchmark. Files are scored recordings or clips; multilingual, code-switching and judge-scored sets are scored per language or label, so no single count applies. Public corpora are named, licensed sets described.
SuiteAudioFilesScored on
Voice agentsScripted caller scenarios voiced by real speakers, recorded in three noise environments, three utterance-length bands and 12 speaker groups12,460Entity error, WER
HealthcarePriMock57 primary-care consultations; licensed clinician dictation; respiratory-clinic recordings479Medical entity error
English short-formCommon Voice; LibriSpeech test-clean and test-other21,588WER
English accentedBritish-, French-, Spanish- and Indian-accented English9,315WER
English long-formTED-LIUM talks; Rev16 podcasts; CALLHOME telephone calls91WER
Judge-scored entitiesShort-form English audio, with entities graded per label by an LLM judge—Entity error (judge)
Non-speechHold music, line noise and silence with no speech2,950Response rate
MultilingualFLEURS and Common Voice test sets in eighteen languages—WER (character error rate for Japanese, Chinese)
Code-switchingFive English↔X pairs derived from read speech, plus a bilingual conversational corpus—WER
DiarizationCALLHOME calls; AMI and NOTSOFAR meetings; DiPCo dinner parties203DER, cpWER

Each test set is marked by how its audio can be obtained: ○ public and open source; ◆ public but licensed or purchasable; ● private, recorded or annotated by us.

How we score

Each metric and how it is computed. The same normalizer and scoring code apply to every system.

WER Word error rate after normalization; lower is better. Character error rate for Japanese, Chinese and Korean.

NEER Normalized entity error rate: the share of entities (names, addresses, numbers, emails, dates, codes, medical and technical terms) transcribed incorrectly.

MER Medical entity error rate over medication and disease terms.

Judge-scored entity error Entity error rate as graded by an LLM judge.

DER / cpWER Diarization error rate, and concatenated minimum-permutation WER.

Non-speech response rate Share of non-speech clips for which the system produces any text.

Streaming harness. Audio is chunked and streamed in real time to each vendor’s endpoint through its documented SDK or WebSocket API, default settings, language fixed. Final transcripts are scored, not partials.

Normalization. One normalizer (case, punctuation, numbers, abbreviations) is applied to references and to every system’s output. Entity metrics are computed on normalized text and, where reported, on formatted text.

Entity error. Entity spans are located in the reference; an entity is an error if any token of its span is wrong in the aligned transcript. The judge-scored variant asks an LLM whether the transcript names the same real-world entity.

Selection rule. Every row is the model’s most recent production run. The overview grid keeps each vendor’s best model per column and names it on hover.

Provenance. Each figure carries its metric and run-date span. Public test sets are named; licensed sets are described.

Why public leaderboards and ours differ

We build for voice agents, clinical documentation, meetings, telephony, and multilingual and code-switched speech, so the test sets cover that range of speakers, accents, languages and acoustic conditions and score what matters in each. Public leaderboards usually ask a narrower question, typically one word-error-rate figure on a few corpora, so a model can rank differently on the two.

Table 22 What each benchmark measures: this report and the public leaderboards reproduced in §9.
What each benchmark measures: this report and the public leaderboards reproduced in §9.
BenchmarkAudioScoresRun byReferences
This reportVoice-agent calls in three noise environments, clinical consultations and dictation, accented and long-form English, eighteen languages, code-switched speech, meetings and telephone calls, non-speech audioEntity error (by class), WER, DER, non-speech response rateAssemblyAI; every vendor on default settingsHuman-verified
Artificial AnalysisAA-AgentTalk (50%), VoxPopuli parliamentary speech (25%), Earnings22 calls (25%)One WER index over the three sets; time to final transcriptArtificial AnalysisPublished with the sets
CovalShort clips under seven recording conditions (clean, accent, far-field, phone codec, reverb…), 7-day windowPooled WER; time to first token; time to final segmentCovalModel-generated
Pipecat / Daily1,000 English samplesSemantic WER; median and P95 time to finalDailyPublished with the set

The Artificial Analysis index is one WER over three sets, half of it AA-AgentTalk, with no entity, noise, medical, multilingual, code-switching or diarization component; its latency runs from a voice-activity detector’s end of speech to the final transcript, which this report does not measure. Coval scores against model-generated references. Neither substitutes for the other or for this report.

Caveats

We ran it. AssemblyAI produced the results in §1–§8. Competitors ran on default settings and may do better tuned. §9 has independent results.

Snapshots differ. Runs span 2026-08-18 to 2026-09-24 (Table 20). A vendor that updated its model within that window may appear on the older version.

Coverage is uneven. Not every system ran every set; means are withheld rather than computed on partial data.

Reference quality. TED-LIUM references omit sponsor read-outs. Some independent leaderboards (Coval) use model-generated references.

Benchmarks

Run your own benchmark

Testing on your own audio? Want to do it right? Read our benchmarking guide in the docs.