New Universal-3.5 Pro is here. Learn more: Async Realtime
Benchmarks

Independently verified: Universal-3.5 Pro on the public speech-to-text benchmarks

Three independent benchmarks — Coval, Daily's Pipecat, and Hugging Face — tested Universal-3.5 Pro on their own audio. Here's what they found.

Independently verified

Universal-3.5 Pro

Top of three independent speech-to-text benchmarks

3.4% Coval WER · #1 of 31 282 ms Pipecat TTFS median Near‑top Open ASR Leaderboard

Written by

David Lange

Published on

22 July 2026

When you're choosing the Voice AI infrastructure your product runs on, you're really making one bet: whether the models underneath are as good as the vendor says. Every provider claims best-in-class accuracy. What buyers actually care about is who can prove it without grading their own homework.

That's why we've been putting more into working with independent benchmarks and the partners who run them. When a third party with no stake in the outcome tests Universal-3.5 Pro against everyone else's models on their own data, the result carries far more weight than anything we could publish ourselves — and lately, across benchmark after benchmark, our models keep landing at or near the top. For a team betting its roadmap on a voice agent or any speech-powered product, that kind of external validation is the clearest signal of model quality there is.

Over the past week alone, three independent benchmarks weighed in: Coval, Daily's open Pipecat STT benchmark, and the Hugging Face Open ASR Leaderboard. Different teams, different audio, same conclusion: leading accuracy at competitive latency, across both realtime and async. Here's what each one found, and why we think third-party validation like this is worth leaning into.

Benchmark 1: Coval

Coval runs a continuously-updated, open-source STT benchmark, and in July 2026 it rebuilt the whole thing to be deliberately harder to game. That's what makes leading it mean something, so it's worth a minute on what Coval actually tests.

Most STT benchmarks lean on the same public datasets — LibriSpeech, Common Voice, TED-LIUM — that model providers also train on. Over time that becomes benchmark contamination: a model can score well by effectively memorizing the test set rather than transcribing real audio well. Coval's fix was a fresh 3,500-sample dataset (up from 50) spanning six conditions clean studio benchmarks skip entirely: clipping, far-field mics, phone-codec compression, reverb, accents, and noise gaps. Those are the conditions real voice agents actually hit — especially over telephony — so no model gets to coast on familiarity with a standard corpus. The benchmark reruns roughly every 30 minutes across 1-, 7-, and 30-day windows, so it reflects the model you'd ship today, not one from three months ago.

Why WER at all? Because in a voice agent the transcript is the input to everything downstream — intent, routing, the action the agent takes. Mishear a medication name or an account number and the whole turn fails. Coval frames WER as the production filter: the first bar a model has to clear before latency or cost even enter the conversation.

Word error rate by model on Coval's STT benchmark, with Universal-3.5 Pro leading at 3.4%.
Word error rate by model on Coval's STT benchmark. Source: Coval, "Word Error Rate 2026"

On that harder dataset, in the 7-day snapshot behind Coval's July 2026 analysis, Universal-3.5 Pro led all 31 models on overall word error rate, at 3.4% — and it topped the field specifically on clipping and far-field audio, two of the toughest real-world conditions. No model wins every condition, but leading overall on a contamination-resistant dataset is exactly the result that predicts real deployment accuracy instead of benchmark familiarity.

Benchmark 2: Daily's Pipecat STT benchmark

Daily maintains Pipecat, the open-source voice-pipeline framework a huge chunk of voice agents are built on. Their STT benchmark runs real agent conversations — 1,000 samples from the smart-turn-data-v3.1 set, not clean read-aloud audio — and scores each provider on two axes that decide whether an agent feels right: semantic WER (only the errors that would change what an LLM understands — punctuation, filler words, and equivalent phrasings are ignored) and TTFS, time to final segment (how long after a speaker stops talking the final transcript lands). Plot the two against each other and you get the Pareto frontier: the services no competitor beats on both.

Daily records a TTFS median of 282 ms at a pooled semantic WER of 1.22% (Source: Daily's Pipecat STT benchmark). Of the seven services on the frontier, only Speechmatics posts a lower semantic WER (1.07%) — and it's the slowest of the group, at 495 ms. Everything faster than us is less accurate: NVIDIA Nemotron (221 ms, 1.95%), Deepgram (247 ms, 1.62%), and Soniox (249–260 ms, 1.27–1.29%) all trade accuracy for their speed. Put simply, the only way to beat Universal-3.5 Pro on accuracy here is to accept more latency.

Pipecat STT Pareto frontier: TTFS median latency vs. pooled semantic WER, with Universal-3.5 Pro on the frontier at 282 ms and 1.22%.
STT Pareto frontier — TTFS median latency vs. accuracy. Source: Daily's Pipecat STT benchmark

The more telling chart is worst-case latency. Daily is explicit that for production voice agents, P95 latency matters more than the median — it's the occasional slow turn, not the typical one, that breaks conversational flow. On the P95 frontier, Universal-3.5 Pro Realtime holds its place at 354 ms and 1.22% semantic WER, while Deepgram drops off the frontier entirely: its median is quick, but its tail isn't. Staying on the frontier at P95 is the harder result, and the one that actually predicts how an agent feels in production.

Pipecat STT Pareto frontier at P95 latency vs. semantic WER, with Universal-3.5 Pro holding the frontier at 354 ms and 1.22%.
STT Pareto frontier — worst-case (P95) latency vs. accuracy. Universal-3.5 Pro holds the frontier at 354 ms and 1.22% semantic WER. Source: Daily's Pipecat STT benchmark

Test it yourself

Test the accuracy on your own audio.

Benchmarks are a starting point, not the finish line. Run your own conversations through Universal-3.5 Pro Realtime and hear how it handles names, numbers, and crosstalk.

Try playground →

Benchmark 3: Hugging Face Open ASR Leaderboard

Pipecat and Coval cover messy, real-world agent audio. The Hugging Face Open ASR Leaderboard is the other end of the spectrum — clean, scripted speech across a standardized set of datasets, run by the community rather than any vendor. Universal-3.5 Pro placed near the top there too.

Doing well on both ends — pristine scripted audio and chaotic agent conversations — is the point. Plenty of models look great on read-aloud datasets and fall apart the moment there's crosstalk, an accent, or a spelled-out order number. Landing near the top of a scripted leaderboard while also holding the Pareto frontier on live conversations is the combination that's hard to fake.

What this means if you're choosing a model

Independent benchmarks are worth more than vendor slides precisely because nobody controls them. Three separate parties — Coval, Daily, and the Open ASR community — measured Universal-3.5 Pro Realtime and async across very different conditions and landed in the same place: the leading overall word error rate on Coval's contamination-resistant dataset, a spot on the Pipecat Pareto frontier that holds even at the P95 tail, and a near-top finish on scripted audio — all at competitive latency. It's also why accuracy has held up as models improve rather than plateauing on paper.

If you're evaluating STT for a voice agent or another use case, don't take our word for it, and don't take anyone's benchmark table at face value — including this one. Run your own audio through it, or wire it straight into the Voice Agent API and hear it in a live conversation. That's the only benchmark that's actually about your product.

Get started

Build on the accuracy leader.

Start streaming with Universal-3.5 Pro Realtime in minutes. Pay-as-you-go pricing, unlimited concurrency, and clear docs — no contract, no commit.

Sign up free →

Frequently asked questions

What is the Pareto frontier in a speech-to-text benchmark?

The Pareto frontier is the set of models where you can't improve one metric without sacrificing another — in Pipecat's benchmark, that means semantic WER versus latency. A model on the frontier can't be beaten on both semantic WER and speed at the same time, so any model behind it is strictly worse on at least one axis. Universal-3.5 Pro Realtime sits on that frontier (Source: Daily's Pipecat STT benchmark).

What is the typical latency of a real-time speech-to-text API?

Real-time speech-to-text APIs typically return final transcripts in a few hundred milliseconds. On Daily's Pipecat benchmark, Universal-3.5 Pro Realtime measured a TTFS (time to final segment) median of 282 ms and a P95 of 354 ms, and it exposes tunable modes — min_latency, balanced, and max_accuracy (Source: Universal-3.5 Pro Realtime launch post). Its end-of-turn detection runs at around 300 ms by reading tonality and pacing rather than silence alone.

Can you build a low-latency voice agent that responds in under 500ms?

Yes. On Daily's Pipecat benchmark, Universal-3.5 Pro Realtime posts a 282 ms median time-to-final (354 ms even at P95), and its end-of-turn detection adds only about 300 ms (Universal-3.5 Pro Realtime launch post), leaving headroom for your LLM and text-to-speech steps. The Voice Agent API bundles speech-to-text, LLM, and TTS through one WebSocket with about 1 second of end-to-end latency (Source: Voice Agent API product page), so you don't have to stitch and tune three providers yourself.

How does Universal-3.5 Pro compare to Deepgram and ElevenLabs on accuracy?

On Coval's STT benchmark, Universal-3.5 Pro led all 31 models on overall word error rate (3.4%), ahead of providers including Deepgram and ElevenLabs (Source: Coval, "Word Error Rate 2026"). On Daily's Pipecat benchmark, which scores semantic WER against latency, it lands on the Pareto frontier at 1.22% pooled semantic WER and a 282 ms TTFS median — only the slower Speechmatics is more accurate, while Deepgram falls off the frontier at P95 (Source: Daily's Pipecat STT benchmark).

How do I start using Universal-3.5 Pro Realtime?

Sign up for a free AssemblyAI account and connect to the streaming WebSocket at wss://streaming.assemblyai.com/v3/ws with "speech_model": "universal-3-5-pro". Pricing is pay-as-you-go at $0.45/hr base with unlimited concurrency and no contract (Source: AssemblyAI pricing / Universal-3.5 Pro Realtime launch post), and you can test it on your own audio in the playground before writing any code.