New Universal-3.5 Pro is here. Learn more: Async Realtime
Audio-to-text API

Convert any audio file to text with one API call

Send a file or URL and get back accurate, formatted text with speaker labels, word-level timestamps, and confidence scores. Universal-3 Pro nails the hard stuff—entities, rare words, and domain terms—out of the box.

Highest accuracy out of the box

Universal-3 Pro captures the hard stuff—entities, rare words, and domain-specific terminology—without tuning.

Any file, one call

Submit audio or video up to 5GB by URL or upload; the SDK uploads, submits, and polls for you across common formats.

Speakers, timestamps, and more

Get speaker diarization, word-level timestamps, and confidence scores in structured JSON, plus optional Speech Understanding add-ons.

Metaview
Ashby
Cluely
Genio
Siro

36%

improvement in close rate

See case study
LiveKit
Earmark
Commure
Dovetail
Fireflies

“The new Universal-3.5 Pro speech model from AssemblyAI is best so far in terms of accuracy, latency, and language switching.”

Retell
CallRail
Apollo.io
ClickUp
Calabrio

80%

increase in customer satisfaction

HeyGen
Granola
Siro
JotPsych
Granola

“Assembly has saved us countless hours managing models, and provided exceptional accuracy.”

Metaview
Ashby
Cluely
Genio
Siro

36%

improvement in close rate

See case study
LiveKit
Earmark
Commure
Dovetail
Fireflies

“The new Universal-3.5 Pro speech model from AssemblyAI is best so far in terms of accuracy, latency, and language switching.”

Retell
CallRail
Apollo.io
ClickUp
Calabrio

80%

increase in customer satisfaction

HeyGen
Granola
Siro
JotPsych
Granola

“Assembly has saved us countless hours managing models, and provided exceptional accuracy.”

Quickstart

From file to transcript in one call

Grab your API key, point the SDK at a file or URL, and it handles upload, submission, and polling for you. You get structured JSON back with everything attached, so you can build and iterate fast.

Start building free

No credit card required

Word error rate

Word error rate is the share of words the model gets wrong against a human reference—the standard measure of transcription accuracy on pre-recorded audio.

Pre-recorded word error rate on English audio.

*Lower is better*
AssemblyAI Universal-3 Pro
4.50%
OpenAI GPT-4o Transcribe
5.34%
Deepgram Nova-3
6.66%
Azure Batch
7.02%

Source: AssemblyAI published benchmarks — assemblyai.com/benchmarks.