New Universal-3.5 Pro is here. Learn more: Async Realtime

Power best in class voice agents

Ultra-fast and ultra-accurate streaming STT built for voice agents. Get 300ms immutable transcripts and intelligent endpointing.

Click to start a voice conversation with our AI support agent

Two solutions

Pick the API that fits your build

Different architectures, different tradeoffs. Both powered by industry-leading speech models.

Recommended

Voice Agent API

Our proprietary voice stack, built on Universal-3.5 Pro, via one WebSocket. Connect, stream audio in, get audio back - we handle the rest.

Get started for free

$4.50 /hr

Speech, LLM, and voice all included

  • Best-in-class voice agents - the preferred way to build with AssemblyAI
  • Customer support agents, AI companions, clinical intake, language learning
  • Teams shipping fast - working agent in an afternoon, no infra to manage
  • Claude Code compatible - paste the docs and build anything

Bring your own stack

Universal-3.5 Pro Realtime

The STT layer for your cascading voice-agent architecture. Works natively with your preferred orchestrator, with Sync Speech-to-Text available for finished utterances.

View integration docs

$0.45 /hr

Transcription only, unlimited concurrent streams

  • Teams already using LiveKit, Pipecat, or Vapi as their orchestration layer
  • Teams running cascading architectures: STT -> LLM -> TTS
  • High-scale deployments where margin and full control matter
  • Complex workflows with RAG, custom tooling, or proprietary LLMs
  • HIPAA, SOC 2 - bring your own compliance infrastructure
  • Sync Speech-to-Text API for teams running their own VAD and turn detection

Choose based on your architecture

Voice Agent API

AssemblyAI's proprietary voice stack

Universal-3.5 Pro Realtime API

Best-in-class STT for your stack

Industry-leading speech models

Unlimited concurrency

Enterprise grade reliability

Session-based pricing

Setup time

Working agent in an afternoon

Minutes to swap STT in an existing stack

Architecture

1 WebSocket · JSON messages · No frameworks required

Cascading (STT → LLM → TTS) — you own the full pipeline

LLM

Managed — update system prompt mid-conversation

Bring your own

Voice (TTS)

Included — select from natural-sounding voices

Bring your own

Pricing

$4.50/hr all-in — no token math across three invoices

$0.45/hr — STT only, unlimited concurrent streams

Integrations

LiveKit, Pipecat, any WebSocket client, Claude Code

LiveKit, Pipecat, custom WebSocket, Twilio SIP

Session resume

30-second reconnect window, context preserved

Via your orchestrator

Integrations

Ready to plug into your voice-agent stack

Pre-built integrations with step-by-step docs enabling quick implementation without disrupting existing workflows.

Common questions

Unlock the power of voice intelligence