New Universal-3.5 Pro is here. Learn more: Async Realtime
Voice Agent API

Ship a voice agent without stitching the stack

One WebSocket gives you speech recognition, turn detection, interruption handling, and tool calling, built on AssemblyAI's real-time Speech Understanding—so you can talk to a working agent today.

One WebSocket

Speech recognition, turn detection, LLM reasoning, and voice over a single connection, so there's no STT, LLM, or TTS stack to wire together.

Natural turns

Built-in turn detection and barge-in handling let users interrupt and be understood the way real conversations work.

Tool calling built in

Hand the agent HTTP tools and it can look up, book, and act mid-conversation without extra orchestration code.

Metaview
Ashby
Cluely
Genio
Siro

36%

improvement in close rate

See case study
LiveKit
Earmark
Commure
Dovetail
Fireflies

“The new Universal-3.5 Pro speech model from AssemblyAI is best so far in terms of accuracy, latency, and language switching.”

Retell
CallRail
Apollo.io
ClickUp
Calabrio

80%

increase in customer satisfaction

HeyGen
Granola
Siro
JotPsych
Granola

“Assembly has saved us countless hours managing models, and provided exceptional accuracy.”

Metaview
Ashby
Cluely
Genio
Siro

36%

improvement in close rate

See case study
LiveKit
Earmark
Commure
Dovetail
Fireflies

“The new Universal-3.5 Pro speech model from AssemblyAI is best so far in terms of accuracy, latency, and language switching.”

Retell
CallRail
Apollo.io
ClickUp
Calabrio

80%

increase in customer satisfaction

HeyGen
Granola
Siro
JotPsych
Granola

“Assembly has saved us countless hours managing models, and provided exceptional accuracy.”

Quickstart

Talk to your first agent minutes after you sign up

Create an agent with one API call, drop it into a small web page, and you're in a live conversation in minutes—echo cancellation is handled for you. Or try a live agent in the playground first.

Start building free

No credit card required

Streaming word error rate

A voice agent is only as good as the words it hears. Streaming word error rate measures real-time transcription accuracy against a human reference—errors here compound through every downstream step.

Streaming word error rate on English audio.

*Lower is better*
AssemblyAI Universal-3.5 Pro Realtime
5.53%
Deepgram Flux
8.87%
Deepgram Nova-3
9.39%
Cartesia
10.10%

Source: AssemblyAI published benchmarks — assemblyai.com/benchmarks.