AssemblyAI's Universal-3.5 Pro Realtime is the only model in Coval's Human Parity Zone
Coval's independent, open-source STT leaderboard shows AssemblyAI's Universal-3.5 Pro Realtime is the only model in the Human Parity Zone — matching or beating humans on both accuracy and speed.



Third-party benchmarks are the closest thing our industry has to a level playing field. Coval's voice-AI benchmarks are exactly that: an open-source harness (Apache-2.0) that runs on a schedule against pinned datasets, scores every provider with the same normalization pipeline, and publishes every number — along with the code to reproduce it.
This week's results are worth sharing. Coval plots every speech-to-text model on two axes — accuracy and latency — and marks out a region it calls the Human Parity Zone. Models in this region match or beat a human on both axes: professional human transcribers achieve 2–4% WER under optimal conditions, and ~200ms is the median gap before a person replies in conversation. Anything inside the zone is at or beyond human performance.
AssemblyAI's Universal-3.5 Pro Realtime is the only model inside it.
What the Human Parity Zone means
Human performance sets a concrete bar on each axis:
- Accuracy: professional human transcribers achieve 2–4% word error rate (WER) under optimal conditions. Universal-3.5 Pro Realtime's average WER over the trailing 7 days is 3.40% — inside the human range, and the lowest of all 29 models on the leaderboard.
- Speed: ~200ms is the median gap before a person replies in conversation. Universal-3.5 Pro Realtime delivers its first speech output under that bar — the response starts before a human listener would even notice a pause.
Every other model on the chart clears one bar or the other. Only Universal-3.5 Pro Realtime clears both. That's the difference between a voice agent that transcribes well or responds naturally, and one that does what a human does: listen accurately and reply on time.
#1 on accuracy
Over the trailing 7-day window (as of July 21, 2026), across roughly 12,000 scored utterances per model:
That's a 19% relative accuracy advantage over the second-place model — and a wider gap still over the models most teams are comparing us against day to day.
#1 on time-to-first-token
Accuracy leaderboards usually come with an asterisk: the most accurate model is often the slowest. Not here. On Coval's latency measurements over the same window, Universal-3.5 Pro Realtime also posted the lowest average time-to-first-token of any model tested — ahead of every model purpose-built for speed, including the streaming-first providers that compete on latency alone.
For voice agents, that combination is the whole game. Latency determines whether a conversation feels natural; accuracy determines whether the agent does the right thing. Coval's data shows you don't have to choose.
That tracks with what builders tell us. Fireflies evaluated the model against the field before moving their real-time pipeline over:
The new Universal 3.5 Pro speech model from Assembly is best so far in terms of accuracy, latency and language switching.
— Foysal Osmany, Software Engineer at Fireflies
LiveKit, which runs voice infrastructure for a large share of the ecosystem, points to the same accuracy edge:
We're excited to make AssemblyAI's Universal-3.5 Pro available on LiveKit Inference. What really stands out is their pace of innovation with Context Carryover — it intelligently applies conversation context to improve transcription accuracy in a way most speech models don't, removing the need for users to predefine key terms.
— David Zhao, Co-founder at LiveKit
Why this benchmark is worth trusting
We like Coval's benchmark because it's hard to game:
- It's open source. The entire harness is on GitHub — datasets, provider configs, metrics code, and the ADRs documenting every methodology decision. Anyone can clone it and re-run the numbers with their own API keys.
- It's not just clean audio. Scoring pools an easy tier (a speaker-balanced subset of LibriSpeech test-clean) with a hard tier: 897 spontaneous, conversational voice-agent clips full of fragments, fillers, and interruptions — the audio real products actually see.
- Normalization is standardized. WER is computed after Whisper's EnglishTextNormalizer — the de facto standard for published WER — applied identically to every provider, so no model wins or loses on formatting quirks.
- Every model runs on identical samples. Within each run, all models are scored on the exact same audio, and datasets are SHA-pinned so results are reproducible.
- It's continuous. This isn't a one-time bake-off with cherry-picked clips. The benchmark runs on a schedule, so the leaderboard reflects live, current API performance — drift and all.
Benchmark it on your own audio
Public benchmarks are a great filter. Your data is the final exam. The fastest way to run it:
- Try the Playground — drop in a real call recording from your product and inspect the transcript quality directly.
- Get a free API key — sign up and start testing with free credits, no credit card required.
- Compare methodically — if you want to reproduce Coval's numbers or design your own test, our guide on how to evaluate speech recognition models walks through datasets, normalization, and the metrics that matter.
A leaderboard can tell you where to look. Your own audio tells you what to ship — and that's the only benchmark that pays your bills.
Frequently asked questions
What is the Human Parity Zone in Coval's speech-to-text benchmark?
The Human Parity Zone is the region of Coval's STT leaderboard where a model matches or beats human performance on both accuracy and speed at the same time. The accuracy bar is the 2–4% word error rate professional transcribers achieve under optimal conditions, and the speed bar is the ~200ms median gap before a person replies in conversation. As of the July 21, 2026 window, AssemblyAI's Universal-3.5 Pro Realtime is the only model inside it.
What word error rate do professional human transcribers achieve?
Professional human transcribers achieve roughly 2–4% word error rate (WER) under optimal listening conditions. That range is why Coval uses it as the accuracy boundary of the Human Parity Zone. Universal-3.5 Pro Realtime's trailing-7-day average WER of 3.40% falls inside that human range.
Is Coval's benchmark independent and reproducible?
Yes. Coval's benchmark is an independently operated, open-source harness (Apache-2.0) published on GitHub, including the datasets, provider configs, metrics code, and methodology decisions. It runs on a schedule against SHA-pinned datasets and scores every provider through the same normalization pipeline, so anyone can clone it and re-run the numbers with their own API keys.
What's the difference between word error rate and time-to-first-token?
Word error rate (WER) measures accuracy — the percentage of words a model transcribes incorrectly — while time-to-first-token measures latency, or how quickly the model returns its first output. For voice agents both matter: latency determines whether a conversation feels natural, and accuracy determines whether the agent does the right thing. Most models lead on one and lag on the other; Universal-3.5 Pro Realtime posted the best score on both in Coval's July window.
Why does accuracy and latency matter so much for voice agents?
A voice agent has to listen accurately and reply on time, just like a person. If latency is high the conversation feels laggy and unnatural; if accuracy is low the agent mishears names, numbers, and commands and takes the wrong action. Winning on only one axis produces an agent that either transcribes well or responds naturally — not both.
How can I test AssemblyAI's accuracy on my own audio?
Start in the AssemblyAI Playground, where you can drop in a real recording and inspect transcript quality with no code or credit card. To test at volume, get a free API key and run your own dataset through the realtime speech-to-text API. If you're evaluating providers for production, our team can help you set up a structured side-by-side comparison.
Benchmark data from benchmarks.coval.ai, retrieved July 21, 2026 (7-day window). Coval's benchmark is independently operated and open source; methodology details are in their repository.
Lorem ipsum dolor sit amet, consectetur adipiscing elit, sed do eiusmod tempor incididunt ut labore et dolore magna aliqua. Ut enim ad minim veniam, quis nostrud exercitation ullamco laboris nisi ut aliquip ex ea commodo consequat. Duis aute irure dolor in reprehenderit in voluptate velit esse cillum dolore eu fugiat nulla pariatur.

