New Universal-3.5 Pro is here. Learn more: Async Realtime
Speech-to-text models

Speech-to-text models built for Voice AI apps

Multilingual speech-to-text with speaker diarization, real-time transcription, and LLM integrations for intelligence out of the box—best-in-class word error rate from $0.15/hr.

Accurate and fast

Best-in-class word error rate and low latency on real-world audio and live streams, across 99 languages.

Real-time diarization

Identify and label speakers on every audio file, streaming or pre-recorded.

Formatted for LLMs

Pipe clean, formatted transcripts straight into LLMs for summarization, topic extraction, and downstream workflows.

Metaview
Ashby
Cluely
Genio
Siro

36%

improvement in close rate

See case study
LiveKit
Earmark
Commure
Dovetail
Fireflies

“The new Universal-3.5 Pro speech model from AssemblyAI is best so far in terms of accuracy, latency, and language switching.”

Retell
CallRail
Apollo.io
ClickUp
Calabrio

80%

increase in customer satisfaction

HeyGen
Granola
Siro
JotPsych
Granola

“Assembly has saved us countless hours managing models, and provided exceptional accuracy.”

Metaview
Ashby
Cluely
Genio
Siro

36%

improvement in close rate

See case study
LiveKit
Earmark
Commure
Dovetail
Fireflies

“The new Universal-3.5 Pro speech model from AssemblyAI is best so far in terms of accuracy, latency, and language switching.”

Retell
CallRail
Apollo.io
ClickUp
Calabrio

80%

increase in customer satisfaction

HeyGen
Granola
Siro
JotPsych
Granola

“Assembly has saved us countless hours managing models, and provided exceptional accuracy.”

Quickstart

Start transcribing minutes after you sign up

Create a free account and transcribe any file with one API call. Test it in the no-code playground, then copy a ready-made request into your app. Usage-based pricing from $0.15/hr, no minimums.

Start building free

No credit card required

Word error rate

Word error rate is the share of words the model gets wrong against a human reference—the standard measure of transcription accuracy on pre-recorded audio.

Pre-recorded word error rate on English audio.

*Lower is better*
AssemblyAI Universal-3 Pro
4.50%
Mistral Voxtral Mini
5.24%
OpenAI GPT-4o Transcribe
5.34%
Deepgram Nova-3
6.66%
Azure Batch
7.02%

Source: AssemblyAI published benchmarks — assemblyai.com/benchmarks.