New Universal-3.6 Pro Realtime is now available Learn more
Insights & Use Cases

How the Voice Agent API pipeline works, from audio in to audio out

A component-by-component walkthrough of the Voice Agent API pipeline for developers who need to understand the internals before trusting it in production.

Abstract green cylinder illustration

Written by

Devon Malloy

Published on

30 September 2026

The Voice Agent API is a managed voice pipeline with seven named stages: noise suppression, streaming speech-to-text, turn detection, LLM inference, text-to-speech, session management, and telephony transport. All seven run behind a single WebSocket at wss://agents.assemblyai.com/v1/ws, billed at a flat $4.50 per agent hour, with roughly one second of end-to-end latency.

"Managed API" is one of the most abused phrases in developer tooling. It usually means: we made some decisions, we're not going to tell you what they are, and when it breaks you get to guess. If you have been burned by a magic API before, that instinct is correct and you should keep it.

So this post does the opposite. Every stage is named, every knob you can turn is listed, every number is sourced, and the parts that genuinely do not exist yet are called out at the end.

The pipeline, named

A voice agent turn is a relay race. Audio comes in, seven stages hand off to each other, audio goes out. Latency is the sum of the handoffs, which is why owning all of them in one process matters more than any single stage being fastest.

1. Voice focus and noise cancellation

The first stage suppresses background noise and isolates the primary speaker before any transcription happens. This ordering is deliberate: speech models are trained on speech, and every unfiltered dog bark, keyboard clatter, or second conversation in the room is an opportunity for the model to transcribe something that was never said.

voice_focus has two profiles. near-field is tuned for headsets and phone handsets, where the speaker is close to the mic and the noise is mostly ambient. far-field is tuned for rooms, kiosks, and drive-thrus, where the speaker is at a distance and competing voices are the dominant problem. Pick the one that matches your deployment, not the one that sounds more capable.

2. Speech-to-text: Universal-3.6 Pro Realtime

The Voice Agent API runs on Universal-3.6 Pro Realtime, our flagship streaming speech model. Transcripts stream out as the user talks rather than arriving after they stop, which is the only way the downstream LLM can start reasoning before the turn ends.

Entity accuracy is the number that actually determines whether a voice agent works. A word error rate of 5% sounds fine until the 5% is the caller's last name, their street, or the last four digits of their account. Here is Universal-3.6 Pro Realtime on AssemblyAI's English voice-agent benchmark, which is scored on 12,460 scripted voice-agent scenarios rather than read audiobooks. Lower is better.

Metric Universal-3.6 Pro Realtime Deepgram Flux EN ElevenLabs Scribe v2 Deepgram Nova-3
Word error rate 5.19% 13.50% 7.78% 8.64%
Entity error rate 14.4% 30.1% 18.5% 26.1%
Names 10.9% 29.0% 14.8% 24.3%
Codes / IDs 10.0% 46.1% 12.0% 27.6%
Phone numbers 2.4% 11.5% 3.4% 4.5%

Independent benchmarks point the same way: on Pipecat's open STT benchmark, it posts a 0.96% pooled semantic word error rate.

Keyterms prompting is included in the Voice Agent API at no extra charge. Pass your SKU list, drug names, or account prefixes and the model biases toward them.

Languages: six. The Voice Agent API supports English, Spanish, French, German, Italian, and Portuguese, with native code-switching across all six inside a single conversation. That is the complete list. A caller can start a sentence in Spanish and finish it in English and the transcript will follow them. If you need coverage beyond those six, that is asynchronous Speech-to-Text territory, not the Voice Agent API — see the supported languages breakdown.

3. Turn detection and interruption handling

End-of-turn detection decides when the user has finished speaking. Silence-threshold VAD gets this wrong constantly, because humans pause mid-thought and then keep going. Universal-3.6 Pro Realtime weighs whether what the user has said sounds complete.

The Voice Agent API layers interruption classification on top. When a user says "uh-huh" or "right" while the agent is talking, that is a backchannel — encouragement, not a request to stop. Treating it as an interruption makes an agent that trails off every three seconds.

On Sesame TurnBench, the full Voice Agent API pipeline produces 55% fewer false positives than Streaming STT alone at the same settings. In practice that is 3.1 wrong interruptions per conversation instead of 6.8. The trade is real and worth naming: it costs about 7 points of recall. The pipeline is tuned to occasionally let a user finish a thought rather than to cut them off, which is the right default for support lookups, scheduling, and qualification.

4. LLM inference through the LLM Gateway

Once the turn closes, the transcript goes to a model through the LLM Gateway. This stage owns system prompt injection, conversation history, and JSON Schema tool calling, and it is the stage you configure most.

Because routing happens inside the same process as transcription, there is no network hop between "the user stopped talking" and "the model started thinking." That is where most of the second of end-to-end latency gets saved.

5. Text-to-speech

The model's reply is synthesized with conversational cadence and streamed back as audio. Voices come from the AssemblyAI voice catalog; pick one in the voices reference. Output volume is a first-class setting, so you can match agent loudness to a drive-thru speaker or a phone handset without post-processing the stream yourself.

6. Session management

Sessions survive brief network loss. Reconnect within 30 seconds and the conversation context is preserved, so a caller walking through a dead zone does not restart the interaction from the top. Echo cancellation and audio format handling live here too.

7. Telephony transport

SIP telephony has shipped. The Voice Agent API plugs into any SIP endpoint, including via Twilio, with AssemblyAI-owned or customer-owned numbers, and it runs under 500ms. DTMF is supported, so callers can enter an account number or extension on the keypad and the agent receives it. Setup is in the Twilio connection guide.

This is the biggest change since this post first ran. Telephony used to be the honest gap in the pipeline. It is not one anymore.

Build on the whole pipeline, not seven vendors

Noise suppression, streaming STT, turn detection, LLM routing, TTS, sessions, and SIP telephony behind one WebSocket. Get an API key and open a session in a few minutes.

Sign up free

Where you steer the pipeline

These are the settings that change agent behavior most, and each one is a single field.

agent_context and Context Carryover

agent_context lets you pass the agent's own question into the speech model, so short or mumbled replies resolve against what was actually asked. Ask "what's your date of birth?" and a mumbled "oh three fourteen eighty-seven" is a date, not noise. Across a benchmark of 20,000 voice agent audio files, agent_context cut WER by 10.2%, with fabrications down 18.3%, hallucinations down 17.2%, place-name entities down 15.5%, short-utterance errors down 13.7%, and name and medical entities each down 9.4%.

Context Carryover is the always-on version: a short rolling conversation memory that biases recognition toward what has already been said in the session. It is on by default and requires no configuration.

"We're excited to make AssemblyAI's Universal-3.5 Pro available on LiveKit Inference. What really stands out is their pace of innovation with Context Carryover — it intelligently applies conversation context to improve transcription accuracy in a way most speech models don't, removing the need for users to predefine key terms."

— David Zhao, Co-founder at LiveKit

voice_focus

Set near-field for headsets and phones, far-field for rooms, kiosks, and drive-thrus. This is the highest-leverage single setting for agents deployed in noisy physical environments, and getting it wrong is a common cause of "the model transcribes the person behind me."

min_latency, balanced, and max_accuracy

Instead of exposing low-level VAD thresholds and forcing you to tune them, the pipeline exposes three modes. min_latency favors responsiveness for short, transactional exchanges. balanced is the default and is the right answer for most agents. max_accuracy trades some speed for correctness, which is what you want when the turn contains an address, a dosage, or a policy number.

Output volume control

Agent output volume is a session setting, not something you handle in your audio layer. Useful when the same agent serves both a web widget and a drive-thru speaker.

Connecting: the actual code

There is no SDK requirement. It is a standard JSON protocol over a WebSocket, which means it works with whatever you already use, including Claude Code.

Opening a session

Connect to wss://agents.assemblyai.com/v1/ws and send one session.update to configure the agent. Note the host: it is agents.assemblyai.com, not api.assemblyai.com.

import asyncio
import json
import websockets

API_KEY = "YOUR_API_KEY"
URL = "wss://agents.assemblyai.com/v1/ws"

SESSION = {
    "type": "session.update",
    "session": {
        "system_prompt": (
            "You are a scheduling assistant for a dental clinic. "
            "Confirm the caller's last name and date of birth before you "
            "look anything up. Keep replies to one or two sentences."
        ),
        # Voice IDs come from the AssemblyAI voice catalog.
        "output": {"voice": "YOUR_VOICE_ID"},
        "input": {"language_codes": ["en"], "voice_focus": "near-field", "transcription_mode": "balanced", "keyterms": ["Okonkwo", "periodontal", "Dr. Reyes"]},
        
        
        
        "tools": [
            {
                "type": "function", "name": "lookup_appointment", "description": "Look up a patient's next scheduled appointment.",
                "parameters": {
                    "type": "object",
                    "properties": {
                        "last_name": {"type": "string"},
                        "date_of_birth": {
                            "type": "string",
                            "description": "YYYY-MM-DD",
                        },
                    },
                    "required": ["last_name", "date_of_birth"],
                },
            }
        ],
    },
}


async def main():
    async with websockets.connect(
        URL, additional_headers={"Authorization": API_KEY}
    ) as ws:
        await ws.send(json.dumps(SESSION))
        async for raw in ws:
            event = json.loads(raw)
            print(event["type"])


asyncio.run(main())

That is the whole setup. System prompt, voice, tools, and the pipeline settings from the previous section, in one message.

A tool call round trip

When the model decides to call a tool, the server sends you a tool.call:

{
  "type": "tool.call",
  "call_id": "tc_01H8XPQ4",
  "name": "lookup_appointment",
  "arguments": {
    "last_name": "Okonkwo",
    "date_of_birth": "1987-03-14"
  }
}

You run the lookup and send back a tool.result with the matching call_id. The agent picks the conversation back up from there:

{
  "type": "tool.result",
  "call_id": "tc_01H8XPQ4",
  "result": {
    "status": "found",
    "appointment": "2026-09-24T15:30:00-04:00",
    "provider": "Dr. Reyes"
  }
}

If your lookup is slow, put the call in hold mode and keep the caller company with a reply.create while you work. Silence is what makes people hang up:

{
  "type": "reply.create",
  "instructions": "Still pulling that up — one moment."
}

Reconfiguring mid-conversation

session.update is not just for setup. Send it again at any point and the change takes effect on the next turn, with no reconnect and no lost context. This is how you gate capabilities behind verification:

{
  "type": "session.update",
  "session": {
    "system_prompt": "The caller is verified. You may now discuss appointment details and offer to reschedule.",
    "input": {"transcription_mode": "max_accuracy"},
    "tools": [
      { "name": "lookup_appointment", "description": "..." },
      { "name": "reschedule_appointment", "description": "..." }
    ]
  }
}

The agent starts the call with one read-only tool and no access to anything sensitive. Once identity checks out, it gets the write tool and a tighter accuracy mode. You never dropped the socket.

Starter repos

Two reference implementations, both dependency-light and both able to deploy to a browser or to a phone number over Twilio SIP:

  • github.com/AssemblyAI/voice-agent-starter-python — Python 3.9+, standard library only
  • github.com/AssemblyAI/voice-agent-starter-js — Node 18+, no dependencies

Clone either one, drop in an API key, and you have a working agent before you have finished reading the rest of this post.

Tool calling in practice

Tool calling is where voice agents actually fail, so it is worth being specific about what changed.

Accuracy went from 69% to 80% on our internal tool-calling dataset. Hallucinated tool arguments went from 2.5% of 3,000 calls to 0% after we added parameter guardrails: arguments can only be inferred from user turns and tool results, never from the model's own generations. A model can no longer invent an account number that sounds plausible, because it is structurally prevented from sourcing one from itself.

HTTP tool calling

Tools can be configured over HTTP and executed on AssemblyAI's side rather than round-tripping to your client. For a straightforward REST lookup, this removes your process from the latency path entirely. Client-side execution is still there when the tool needs your credentials or your private network. Both paths are covered in the tool calling docs.

Tool-call parameter hints

Parameter hints tell the pipeline what kind of entity is being spoken when a tool argument is being collected. That improves both recognition accuracy on the entity and turn detection around it, because the pipeline knows a caller reciting a 10-digit number is not finished at the first pause.

Tool-call hold mode

Hold mode keeps the session responsive while a slow tool runs. Combine it with manual reply.create status updates, as in the code above, and a four-second database query stops sounding like a dropped call.

Managing agents outside the socket

Agent Management API

Agent configurations do not have to be assembled in your client on every connection. POST a config to the Agent Management API, store it, and reference it. Prompt and tool changes become a deploy rather than a code change across every service that opens a session.

Webhooks

Webhooks deliver session and telephone-call lifecycle events to your backend. This is how you drive CRM writes, post-call summaries, or billing off agent activity without polling or scraping your own logs.

Where you can customize

You control The API manages
System prompt and all LLM instructions Speech-to-text inference and streaming
Tool definitions (JSON Schema) and parameter hints End-of-turn detection model execution
Voice selection and output volume Interruption classification (backchannel vs. intent to speak)
Pipeline mode: min_latency, balanced, max_accuracy LLM routing and request formation
Keyterms, language, agent_context, voice_focus TTS synthesis, echo cancellation, audio format handling

System prompt, tool definitions, voice, keyterms, and pipeline settings can all be changed mid-conversation with a session.update. You do not reconnect and you do not lose context.

Observability: one place to look

Every event in a turn lands in a single conversation view with timestamps: speech input, transcript, LLM request, LLM response, TTS generation, and audio output. Speaker diarization data is included for multi-party calls.

The point is not that the dashboard is pretty. The point is that when an agent gives a wrong answer, you can see in one place whether the transcript was wrong, the model reasoned badly, or the tool returned bad data. Debugging that across a separate STT vendor, an LLM provider, and a TTS vendor means correlating three sets of logs on three different clocks, and it is where teams lose days.

Hear the pipeline before you wire it up

Test transcription accuracy, code-switching, and entity recognition on your own audio in the browser — no key, no setup.

Try playground

What the API doesn't do

An honest list is more useful than a feature grid, so here is what is genuinely still out of scope.

You cannot point it at your own LLM endpoint. Model selection happens through the LLM Gateway catalog. If you need to run a fine-tuned model on your own inference infrastructure inside the turn loop, this is not the product for that today.

No voice cloning. You choose from the AssemblyAI voice catalog. There is no path to upload a reference sample and synthesize a specific person's voice.

Six languages, and only six. English, Spanish, French, German, Italian, and Portuguese. Our asynchronous Speech-to-Text products cover far more, but the Voice Agent API does not, and building a Japanese-language agent on it is not currently possible.

BAA-backed deployments do not cover the Voice Agent API yet. See the healthcare section below for the current path.

It is a pipeline, not an orchestration framework. If your application needs per-stage control the API does not expose — custom logic between transcript and LLM, or a bespoke turn-taking policy — assemble it yourself from Streaming Speech-to-Text and the LLM Gateway. That path stays fully supported and is the right call for genuinely unusual architectures.

Telephony is no longer on this list. SIP shipped, it runs under 500ms, and DTMF works.

Healthcare and PHI workloads

AssemblyAI is considered a business associate under HIPAA, and we offer a standard Business Associate Addendum (BAA) that is required under HIPAA to ensure that AAI appropriately safeguards PHI. AssemblyAI enables covered entities and their business associates subject to HIPAA to use the AssemblyAI services to process protected health information (PHI). Details are on our BAA legal page.

Scoping matters here, so be precise about it: BAA-backed deployments run through the Speech-to-Text and Streaming Speech-to-Text products. For healthcare use cases on the Voice Agent API specifically, contact sales rather than assuming coverage. AssemblyAI also holds SOC 2 Type 2, ISO 27001:2022, and PCI DSS v4.0.

What $4.50/hr is actually buying you

$4.50 per agent hour ($0.075/min) is flat and all-inclusive: noise suppression, Universal-3.6 Pro Realtime, keyterms prompting, turn detection, interruption classification, LLM inference, TTS, session management, telephony, observability, and the tool-calling machinery.

Priced separately, those components approach that number before you add anything else. What the flat rate really removes is the work around them: the integration between five vendors, the retry logic at each boundary, the three sets of logs, the five contracts, and the on-call rotation for a stack where any vendor's incident is your incident. One bill, one set of logs, one place to file a ticket. Current pricing for every product is on our pricing page.

The pipeline is tuned for structured, transactional work: support lookups, scheduling, ordering, qualification, and triage. That is where the turn-detection trade-offs, the entity accuracy, and the tool-calling guardrails all point.

Talk through your voice agent architecture

Volume pricing, SIP deployment, EU data residency, or healthcare scoping — walk through the pipeline with an engineer who has shipped it.

Talk to AI expert

Frequently asked questions

What is the Voice Agent API?

The AssemblyAI Voice Agent API is a managed voice pipeline that bundles noise suppression, streaming speech-to-text, turn detection, LLM inference, text-to-speech, session management, and SIP telephony behind a single WebSocket connection at wss://agents.assemblyai.com/v1/ws. It is built on Universal-3.6 Pro Realtime, runs at roughly one second of end-to-end latency, and costs a flat $4.50 per agent hour. It uses a standard JSON protocol, so no SDK is required.

How does the Voice Agent API differ from the Streaming STT API?

Streaming Speech-to-Text gives you transcription only — you supply and integrate the LLM, TTS, turn-taking logic, and telephony yourself. The Voice Agent API gives you the entire conversational pipeline through one connection at $4.50 per agent hour, including turn detection, interruption classification, tool calling, and SIP. Choose Streaming STT when you need per-stage control the managed pipeline does not expose; choose the Voice Agent API when you want the pipeline to be someone else's operational problem.

What speech model does the Voice Agent API use?

The Voice Agent API runs on Universal-3.6 Pro Realtime, AssemblyAI's flagship streaming speech model. On AssemblyAI's English voice-agent benchmark it records a 5.19% word error rate and a 14.4% entity error rate, compared with 13.50%/30.1% for Deepgram Flux EN, 7.78%/18.5% for ElevenLabs Scribe v2, and 8.64%/26.1% for Deepgram Nova-3. Keyterms prompting is included at no extra charge.

Can I change the system prompt or tools mid-conversation?

Yes. Send a session.update message at any point during a live session and the new system prompt, tool definitions, voice, keyterms, or pipeline mode take effect on the next turn. You do not reconnect and you do not lose conversation context, which is how teams gate write-access tools behind an identity check inside a single call.

Which languages does the Voice Agent API support?

The Voice Agent API supports six languages: English, Spanish, French, German, Italian, and Portuguese. It handles native code-switching across all six, so a caller can change language mid-sentence and the transcript follows. Broader language coverage is available through AssemblyAI's asynchronous Speech-to-Text products, not the Voice Agent API.

Does the Voice Agent API support phone calls?

Yes. SIP telephony has shipped and connects to any SIP endpoint, including through Twilio, using AssemblyAI-owned or customer-owned phone numbers. Telephony latency runs under 500ms, and DTMF keypad entry is supported so callers can enter account numbers or extensions. Both starter repos can deploy an agent to a phone number over Twilio SIP.

How accurate is the Voice Agent API at tool calling?

Tool-calling accuracy improved from 69% to 80% on AssemblyAI's internal evaluation dataset. Hallucinated tool arguments went from 2.5% of 3,000 calls to 0% after parameter guardrails were added: tool arguments can only be inferred from user turns and tool results, never from the model's own generations. Tools can execute client-side or over HTTP on AssemblyAI's side.

Can I use the Voice Agent API for healthcare or PHI workloads?

AssemblyAI is considered a business associate under HIPAA, and we offer a standard Business Associate Addendum (BAA) that is required under HIPAA to ensure that AAI appropriately safeguards PHI. Today those BAA-backed deployments run through the Speech-to-Text and Streaming Speech-to-Text products. For healthcare or PHI use cases on the Voice Agent API specifically, contact AssemblyAI sales to confirm scope before you build.