New Universal-3.5 Pro is here. Learn more: Async Realtime
Speech-to-Speech

Real-time speech-to-speech without the pipeline

Speak to your app and have it speak back. AssemblyAI's Voice Agent API manages the full real-time loop—speech recognition, LLM reasoning, and voice—over a single WebSocket, so you don't stitch together separate providers.

Audio in, audio out

Stream a user's voice in and get a spoken reply back over one connection, with interruptions handled for you.

Real-time by default

Turn detection and low-latency streaming keep the back-and-forth feeling like a conversation, not a walkie-talkie.

Accuracy at the front

The loop is built on AssemblyAI's real-time speech recognition, so the words feeding your model are right the first time.

Metaview
Ashby
Cluely
Genio
Siro

36%

improvement in close rate

See case study
LiveKit
Earmark
Commure
Dovetail
Fireflies

“The new Universal-3.5 Pro speech model from AssemblyAI is best so far in terms of accuracy, latency, and language switching.”

Retell
CallRail
Apollo.io
ClickUp
Calabrio

80%

increase in customer satisfaction

HeyGen
Granola
Siro
JotPsych
Granola

“Assembly has saved us countless hours managing models, and provided exceptional accuracy.”

Metaview
Ashby
Cluely
Genio
Siro

36%

improvement in close rate

See case study
LiveKit
Earmark
Commure
Dovetail
Fireflies

“The new Universal-3.5 Pro speech model from AssemblyAI is best so far in terms of accuracy, latency, and language switching.”

Retell
CallRail
Apollo.io
ClickUp
Calabrio

80%

increase in customer satisfaction

HeyGen
Granola
Siro
JotPsych
Granola

“Assembly has saved us countless hours managing models, and provided exceptional accuracy.”

Quickstart

From zero to a talking app, fast

Create an agent with a single API call and connect over one WebSocket—speech recognition, reasoning, and voice are already wired together, so your first real-time exchange happens in minutes. Try it in the playground with no code first.

Start building free

No credit card required

Streaming word error rate

A speech-to-speech loop is only as good as the words it hears. Streaming word error rate measures real-time transcription accuracy at the front of the loop against a human reference—errors here compound through reasoning and the spoken reply.

Streaming word error rate on English audio.

*Lower is better*
AssemblyAI Universal-3.5 Pro Realtime
5.53%
Deepgram Flux
8.87%
Deepgram Nova-3
9.39%
Cartesia
10.10%

Source: AssemblyAI published benchmarks — assemblyai.com/benchmarks.