AssemblyAI Universal-3-Pro vs ElevenLabs Scribe v2 Compared
ElevenLabs Scribe v2 vs. AssemblyAI Universal-3.5 Pro: a head-to-head on accuracy, code-switching, entity errors, diarization, latency, and pricing to help you pick the right speech-to-text model.



If you're choosing a speech-to-text engine in 2026, ElevenLabs Scribe v2 and AssemblyAI Universal-3.5 Pro are almost certainly on your shortlist. They're two of the most accurate transcription models on the market, and they come at the problem from different directions: Scribe v2 is the ASR model inside ElevenLabs' broader voice-generation platform, and Universal-3.5 Pro is AssemblyAI's flagship transcription-and-understanding model, purpose-built for teams that put audio at the center of their product.
I lead Voice AI at AssemblyAI, so I'm not going to pretend this is a neutral referee's writeup. But I care more about you making the right call than about winning every row in a table, so I'll show you the benchmarks, tell you where Scribe v2 is genuinely strong, and be explicit about what's public data versus what's ours. Let's get into it.
The quick verdict
Here's the short version if you're skimming. On the accuracy dimensions that actually break production systems—word error rate on messy code-switched audio, entity accuracy in real voice-agent conversations, and speaker diarization quality—Universal-3.5 Pro comes out ahead in our head-to-head testing. Scribe v2 is a strong model, particularly for batch subtitling and captioning at scale, and it leads on some public benchmarks. If you're already all-in on the ElevenLabs stack for voice generation, Scribe is a reasonable default. If transcription accuracy, diarization, and understanding are the product, Universal-3.5 Pro is the one I'd build on.
At-a-glance comparison
A quick note before the detailed breakdown: these head-to-head numbers come from AssemblyAI's own benchmark suite, run on real code-switched audio, real voice-agent conversations, and diarization test sets. I'll flag where public benchmarks tell a different story so you can weigh both.
Accuracy: the benchmarks that matter
Most comparison posts wave at accuracy and move on. That's a mistake, because accuracy isn't one number—it's several, and the ones that matter depend on your audio. Here are the three that decide most production deployments.
Code-switching word error rate
Real-world audio isn't clean single-language speech. People switch between English and Spanish mid-sentence, drop in product names, and talk over each other. On our normalized code-switching benchmark (average WER, lower is better), Universal-3.5 Pro scores 7.69 versus Scribe v2 at 8.77. For reference, the previous flagship Universal-3 Pro scored 9.07 and Deepgram Nova-3 ML scored 12.22, so the whole field has tightened—but 3.5 Pro leads it.
Worth knowing: an independent ServiceNow and Hugging Face code-switching ASR benchmark from June 2026 ranked ElevenLabs Scribe v2, Gemini 3 Flash, and AssemblyAI's Universal-3 line at the top of the field across WER, SWER, and AER. Both models are genuinely elite here. Our internal edge on 3.5 Pro is real, but this is a close, respectable race.
Voice-agent accuracy and entity errors
This is where the gap widens. On the Pipecat open STT benchmark—real agent conversations, not read-aloud scripts—Universal-3.5 Pro Realtime posts 6.99% WER versus Scribe v2 at 9.76%. For context, Google Chirp3 lands at 9.04% and Deepgram Flux at 15.58%.
But the WER number isn't the headline. The entity error rate is. Universal-3.5 Pro Realtime drops 15.31% of entities—names, places, phone numbers—while Scribe v2 drops 39.70%. That's roughly 2.5x more errors on exactly the tokens that matter in a real conversation. If your voice agent mishears a confirmation number or a customer's name, the transcript's overall WER being "good" is cold comfort. Breaking it down further, Universal-3.5 Pro Realtime hits 16.92% error on names, 6.28% on places, and 3.55% on phone numbers. For voice agents and contact centers, entity accuracy is often the whole ballgame.
Diarization quality
Knowing who said what is as important as knowing what was said, and it's where a lot of transcription products quietly fall down. Universal-3.5 Pro is the most accurate diarization we've ever shipped, optimized specifically for cpWER (concatenated minimum-permutation word error rate, which penalizes both transcription and speaker-attribution mistakes). On our diarization benchmark it scores 30.17 versus Scribe v2 at 35.26 and Deepgram Nova-3 EN at 37.92.
This isn't academic. Speaker-label quality is repeatedly the deciding factor in real evaluations of meeting-intelligence and contact-center products—it's the difference between a usable transcript and one a rep has to re-listen to. If your use case involves multi-speaker calls, weight this row heavily.
Public vs private benchmark data—read this before you decide
Here's the honest caveat, because you'll see conflicting numbers if you shop around. On some public, open-source datasets, Scribe v2 leads. On certain private evaluation sets we've run, it can win specific WER slices too. Models that are tuned hard against well-known public benchmarks sometimes look stronger there than they do on the messy, private, real-world data you'll actually feed them. Our numbers above come from benchmarks built to mirror production audio—code-switched speech, live agent calls, real multi-speaker recordings. When you run your own evaluation (and you should), test on your audio, not a leaderboard. A model that tops a public chart but stumbles on your calls is worse than one that's a hair behind on paper and steady in production. For how we think about this, see our writeups on evaluating speech recognition models and why WER alone is broken.
Entity detection and redaction
Beyond raw entity accuracy, both platforms detect and redact sensitive information. AssemblyAI's Speech Understanding layer handles PII redaction, entity detection, and content moderation as part of the same API call, and you can feed transcripts straight into the LLM Gateway to run summarization or custom extraction against GPT, Claude, or Gemini without stitching providers together. If your pipeline needs to catch and mask names and phone numbers reliably, the 15.31% vs 39.70% entity gap above is the number to sit with—you can't redact what the model didn't hear.
Language support
This is a spot where the old version of this comparison undersold us, so let me correct the record. Universal-3.5 Pro does native code-switching across 18 languages—meaning it handles speakers who mix languages within a single utterance, not just one language per file. For broader coverage, Universal-2 supports 99+ languages as a fallback.
Scribe v2 advertises 90+ languages, and if your only requirement is "transcribe this file in language X," that breadth is real. But raw language count and native code-switching are different capabilities. If your audio has speakers moving between English and Spanish, or Hindi and English, mid-sentence—which is most real-world multilingual audio—the 18-language native code-switching support plus the 7.69 code-switching WER is the combination that matters.
Pricing
Universal-3.5 Pro async runs $0.21/hr—flat, transparent, no per-seat licensing. (An earlier version of this post listed $0.20 in one place; the correct flagship async price is $0.21/hr.) The realtime flagship, Universal-3.5 Pro Realtime, is $0.45/hr base. If you need speaker diarization, entity detection, and redaction, those are part of the platform rather than a separate SKU to negotiate. You can see the full breakdown on our pricing page.
ElevenLabs prices Scribe within its broader platform plans, so the effective cost depends on what tier you're on and what else you're using. If you're already paying for ElevenLabs for voice generation, adding Scribe may look incrementally cheap. If transcription is your primary spend, compare the all-in per-hour cost honestly, including the features you'd otherwise bolt on.
Latency and realtime
For streaming and voice agents, latency is a feature. Universal-3.5 Pro Realtime delivers roughly ~300ms end-of-turn detection and is the foundation under our realtime flagship and Voice Agent API, with live diarization, voice_focus, and configurable min-latency, balanced, and max-accuracy modes. Scribe v2 Realtime claims sub-150ms WebSocket latency, which on paper is faster to first token.
Here's the nuance: raw latency-to-token and end-of-turn detection aren't the same thing. A model that emits tokens fast but is jumpy about deciding when a speaker is done can feel worse in a real agent, because it interrupts. Our realtime stack is tuned for natural turn-taking, and the 6.99% voice-agent WER and 15.31% entity error above were measured on that realtime model in real conversations. Early adopters building on it include Retell, LiveKit, and Fireflies. If you're evaluating for a voice agent specifically, dig into our breakdown of streaming speech-to-text and streaming speaker diarization.
Developer experience
Both are developer-friendly APIs, but the design goals differ. AssemblyAI is built to be the single transcription-and-understanding layer you don't outgrow: one API for async and streaming, an async speech-to-text endpoint where you can omit the model parameter to auto-upgrade to the current flagship, and Speech Understanding and LLM Gateway on the same account. On the async quality side, Metaview reported a roughly 47% drop in low-confidence tokens after moving to Universal-3.5 Pro async—the kind of improvement you feel downstream in summaries and analytics. ElevenLabs' developer experience is strong too, but it's organized around a voice-generation platform, so transcription is one capability among many rather than the core product.
When to choose AssemblyAI Universal-3.5 Pro
- Transcription accuracy, diarization, or understanding is central to your product, not a side feature.
- You're building voice agents or contact-center analytics where entity accuracy (names, numbers, confirmations) is critical—that 15.31% vs 39.70% gap is your gap.
- You have multi-speaker audio and need best-in-class diarization (cpWER 30.17).
- Your audio is multilingual with real code-switching, not just one language per file.
- You want transcription, redaction, and LLM-powered analysis on one API without stitching vendors together.
When to consider ElevenLabs Scribe v2
- You're already standardized on ElevenLabs for voice generation and want one vendor across TTS and STT.
- Your primary workload is batch subtitling, captioning, or media transcription at scale, which is exactly what Scribe v2 was positioned for at its January 2026 launch.
- Raw sub-150ms first-token latency matters more to your architecture than end-of-turn quality.
- Your evaluation on your own audio shows Scribe leading on the specific slices you care about—in which case, trust your data over anyone's blog.
The verdict
Universal-3.5 Pro wins the accuracy dimensions that break production systems: code-switching WER (7.69 vs 8.77), voice-agent WER (6.99% vs 9.76%), entity errors (15.31% vs 39.70%), and diarization cpWER (30.17 vs 35.26). Scribe v2 is a genuinely strong model that leads on parts of the public benchmark landscape and fits naturally if you live inside the ElevenLabs ecosystem. My recommendation: if audio understanding is your product, build on Universal-3.5 Pro—and prove it to yourself by running both on your own data before you commit. The insight I'd leave you with is that the "which is more accurate" question has no single answer; the real question is which model degrades gracefully on your hardest audio, and that's something only your evaluation can tell you.
Frequently asked questions
Is ElevenLabs Scribe v2 or AssemblyAI more accurate?
On AssemblyAI's head-to-head benchmarks, Universal-3.5 Pro leads on code-switching WER (7.69 vs 8.77), voice-agent WER (6.99% vs 9.76%), entity error rate (15.31% vs 39.70%), and diarization cpWER (30.17 vs 35.26). Scribe v2 leads on some public datasets. The accurate answer is "it depends on your audio"—so test both on your own data.
What is the difference between Scribe v2 and Scribe v2 Realtime?
Scribe v2 is ElevenLabs' async model, positioned for batch transcription, subtitling, and captioning at scale. Scribe v2 Realtime is the streaming variant, which claims sub-150ms WebSocket latency for live use cases like voice agents and captions.
Is there an ElevenLabs Scribe alternative?
Yes—AssemblyAI Universal-3.5 Pro is a purpose-built transcription-and-understanding alternative with leading diarization, native code-switching across 18 languages, and Speech Understanding plus LLM Gateway on the same API. You can try it free and benchmark it against Scribe on your own audio.
How much does AssemblyAI Universal-3.5 Pro cost?
Async Universal-3.5 Pro is $0.21/hr and the realtime flagship (Universal-3.5 Pro Realtime) is $0.45/hr base, with diarization, entity detection, and redaction included in the platform rather than sold separately. See the pricing page for details.
How many languages does Universal-3.5 Pro support?
Universal-3.5 Pro does native code-switching across 18 languages—handling speakers who mix languages within a single utterance. For broader single-language coverage, Universal-2 supports 99+ languages as a fallback.
Which is better for voice agents?
For voice agents, Universal-3.5 Pro Realtime is the stronger choice in our testing: 6.99% voice-agent WER versus Scribe v2's 9.76%, and a 15.31% entity error rate versus 39.70%. Entity accuracy on names, numbers, and confirmations is usually what makes or breaks an agent, and that's the widest gap between the two.
Try it yourself
Try AssemblyAI free and run Universal-3.5 Pro against Scribe v2 on your own audio.
See the full benchmarks for code-switching WER, voice-agent WER, and diarization cpWER.
Read the Universal-3.5 Pro deep dive to understand what changed in the async flagship.
Lorem ipsum dolor sit amet, consectetur adipiscing elit, sed do eiusmod tempor incididunt ut labore et dolore magna aliqua. Ut enim ad minim veniam, quis nostrud exercitation ullamco laboris nisi ut aliquip ex ea commodo consequat. Duis aute irure dolor in reprehenderit in voluptate velit esse cillum dolore eu fugiat nulla pariatur.




