New Universal-3.5 Pro is here. Learn more: Async Realtime
Real-time speech-to-text

Turn audio into text in real time

Stream audio over a WebSocket and get formatted, immutable transcripts back with sub-300ms latency—accurate enough for production voice agents. Open a connection and start transcribing in a few lines of Python or JavaScript.

Sub-300ms latency

Universal Streaming returns formatted, immutable transcripts fast enough to drive live voice agents and agent-assist.

Built for live conversation

Native turn detection, speaker labels, and keyterm prompting handle the hard parts of streaming audio out of the box.

Real-time multilingual

Stream English plus Spanish, German, French, Italian, and Portuguese, with native code-switching mid-utterance.

Metaview
Ashby
Cluely
Genio
Siro

36%

improvement in close rate

See case study
LiveKit
Earmark
Commure
Dovetail
Fireflies

“The new Universal-3.5 Pro speech model from AssemblyAI is best so far in terms of accuracy, latency, and language switching.”

Retell
CallRail
Apollo.io
ClickUp
Calabrio

80%

increase in customer satisfaction

HeyGen
Granola
Siro
JotPsych
Granola

“Assembly has saved us countless hours managing models, and provided exceptional accuracy.”

Metaview
Ashby
Cluely
Genio
Siro

36%

improvement in close rate

See case study
LiveKit
Earmark
Commure
Dovetail
Fireflies

“The new Universal-3.5 Pro speech model from AssemblyAI is best so far in terms of accuracy, latency, and language switching.”

Retell
CallRail
Apollo.io
ClickUp
Calabrio

80%

increase in customer satisfaction

HeyGen
Granola
Siro
JotPsych
Granola

“Assembly has saved us countless hours managing models, and provided exceptional accuracy.”

Quickstart

Open a stream minutes after you sign up

Grab your API key and connect to the streaming WebSocket with the Python or JavaScript SDK—no infrastructure to stand up. Handle partial and finalized turns with a single event handler and start building.

Start building free

No credit card required

Streaming word error rate

Streaming word error rate measures real-time transcription accuracy against a human reference—the core signal for anything that acts on live audio.

Streaming word error rate on English audio.

*Lower is better*
AssemblyAI Universal-3.5 Pro Realtime
5.53%
Deepgram Flux
8.87%
Deepgram Nova-3
9.39%
Cartesia
10.10%

Source: AssemblyAI published benchmarks — assemblyai.com/benchmarks.