New Universal-3.5 Pro is here. Learn more: Async Realtime
Python SDK

Add speech-to-text to Python in a few lines

Install the SDK, drop in your API key, and transcribe a file with one call—the SDK handles upload, submission, and polling for you. Start free, no credit card.

One pip install

A single pip install adds production speech-to-text to your Python project—no other dependencies to set up.

One call, fully managed

Transcriber().transcribe() uploads, submits, and polls in a single line, so there's no queue plumbing to write.

Add features inline

Flip on speaker labels, language detection, or entity detection by setting one option on the same request.

Metaview
Ashby
Cluely
Genio
Siro

36%

improvement in close rate

See case study
LiveKit
Earmark
Commure
Dovetail
Fireflies

“The new Universal-3.5 Pro speech model from AssemblyAI is best so far in terms of accuracy, latency, and language switching.”

Retell
CallRail
Apollo.io
ClickUp
Calabrio

80%

increase in customer satisfaction

HeyGen
Granola
Siro
JotPsych
Granola

“Assembly has saved us countless hours managing models, and provided exceptional accuracy.”

Metaview
Ashby
Cluely
Genio
Siro

36%

improvement in close rate

See case study
LiveKit
Earmark
Commure
Dovetail
Fireflies

“The new Universal-3.5 Pro speech model from AssemblyAI is best so far in terms of accuracy, latency, and language switching.”

Retell
CallRail
Apollo.io
ClickUp
Calabrio

80%

increase in customer satisfaction

HeyGen
Granola
Siro
JotPsych
Granola

“Assembly has saved us countless hours managing models, and provided exceptional accuracy.”

Quickstart

From install to transcript in minutes

Grab a free API key, run pip install assemblyai, and paste a few lines into a script. Run it and your transcript prints to the terminal—no infrastructure, no async queue to manage.

Start building free

No credit card required

Word error rate

Word error rate is the share of words the model gets wrong against a human reference—the standard measure of transcription accuracy on pre-recorded audio.

Pre-recorded word error rate on English audio.

*Lower is better*
AssemblyAI Universal-3 Pro
4.50%
OpenAI GPT-4o Transcribe
5.34%
Deepgram Nova-3
6.66%
Azure Batch
7.02%

Source: AssemblyAI published benchmarks — assemblyai.com/benchmarks.