Insights & Use Cases
September 30, 2026

How to build a voice agent with Twilio and AssemblyAI

Build an inbound phone voice agent that bridges Twilio Media Streams into AssemblyAI Universal-3 Pro Streaming, GPT-4o with tool calling, and ElevenLabs TTS — all under an 800ms turn budget. Full Python code, deployment guide, and forkable repo included.

Kelsey Foster
, 
Growth
Reviewed by
No items found.
Abstract green mobius illustration
Table of contents

There are two ways to put an AI voice agent on a Twilio phone number, and picking the right one is most of the work. The short path is managed SIP telephony: point a Twilio SIP endpoint at AssemblyAI's Voice Agent API, deploy a starter repo, and you have a working phone agent in well under 500ms of round-trip latency without writing a bridge. The long path is a bridge you own — Twilio Media Streams into Universal-3.6 Pro Realtime, your own LLM, your own TTS — which costs more engineering time but gives you full control over every hop.

This guide covers both. Path A gets you to a ringing phone number fastest. Path B is the complete FastAPI tutorial, updated, for teams who need to own the pipeline.

Which path should you take?

Path A — managed SIP telephony Path B — DIY FastAPI bridge
What you connect A Twilio SIP endpoint to the Voice Agent API Twilio Media Streams to Streaming STT, then your LLM and TTS
What you write A system prompt and your tool handlers A WebSocket server, audio plumbing, turn orchestration, barge-in, tools
Latency Under 500ms for the SIP leg; ~1 second end-to-end Yours to tune — and yours to blow
Billing Flat $4.50/hr covering STT + LLM + TTS, plus Twilio Separate STT, LLM, TTS and Twilio line items
Languages Six, with native code-switching 32 for transcription; LLM and TTS coverage is your choice
Pick it when You want a phone agent shipped this week and the standard pipeline is fine You need a specific LLM, a specific voice, custom turn logic, or an existing pipeline you can't replace

Most teams building a support line, a scheduling agent, or a qualification bot should start with Path A and only drop to Path B when they hit something the managed path won't do. Teams who already run a Pipecat or LiveKit pipeline, or who have a fine-tuned model they must route through, belong in Path B.

Why AssemblyAI for phone audio specifically

Phone audio is 8kHz mulaw. That is a narrow, lossy, compressed format built in the 1970s, and most speech models were trained on 16kHz-and-up wideband audio. The usual workaround is to resample on the way in, which costs latency and does not recover the information the codec already threw away.

Streaming Speech-to-Text accepts 8kHz mulaw natively in both directions. You forward Twilio's frames through unchanged and stream mulaw back out. No resampling stage, no format negotiation, no extra buffer.

The accuracy argument matters more on a phone line than anywhere else, because phone agents spend their whole lives collecting things that have to be exactly right: order numbers, phone numbers, names, addresses, confirmation codes. Here is Universal-3.6 Pro Realtime on AssemblyAI's English voice-agent benchmark, 12,460 scripted voice-agent scenarios built around the entities agents have to capture (lower is better):

Metric Universal-3.6 Pro Realtime Deepgram Flux EN ElevenLabs Scribe v2 Deepgram Nova-3
Word error rate 5.19% 13.50% 7.78% 8.64%
Entity error rate 14.4% 30.1% 18.5% 26.1%
Names 10.9% 29.0% 14.8% 24.3%
Codes / IDs 10.0% 46.1% 12.0% 27.6%
Phone numbers 2.4% 11.5% 3.4% 4.5%

Read the phone numbers row first. A 2.4% error rate against Deepgram Flux EN's 11.5% is the difference between an agent that reads a callback number back correctly and one that asks the caller to repeat it. The names row (10.9% against 29.0%) is the difference between an agent that can look up an account and one that transfers to a human. On Pipecat's open STT benchmark, which measures real agent conversations, it independently posts a 0.96% pooled semantic WER. More comparisons live on the benchmarks page.

This is the category CallRail works in — call tracking and phone-based marketing attribution, where every transcript feeds an attribution decision:

"The capabilities AssemblyAI enables us to build help businesses market confidently while saving time and money. It's powerful, almost magical to see it work."

— Ryan Johnson, Chief Product Officer, CallRail

Put a voice agent on a phone number

Get an API key and start streaming 8kHz mulaw from Twilio in a few minutes. No resampling, no format conversion, and free credit to test with.

Sign up free

Path A: managed SIP telephony (the short path)

SIP telephony is shipped. The Voice Agent API plugs into any SIP endpoint, which includes Twilio's, and the SIP leg runs under 500ms. You can bring your own number or use an AssemblyAI-owned one, and DTMF is supported, so keypad entry works out of the box.

The fastest way in is a starter repo. Both of these support deploying to a browser or to a phone number over Twilio SIP:

The shape of the work is:

  1. Clone a starter and add your API key.
  2. Write the system prompt and register your tools. Keep the scope structured and transactional — order lookups, appointment scheduling, triage, qualification. Those are the use cases phone agents are good at.
  3. Create a SIP trunk or SIP domain in Twilio and point it at the agent endpoint.
  4. Route your Twilio number to that SIP endpoint.
  5. Call the number.

That is the whole build. There is no bridge to write, no audio format to negotiate, no turn orchestration to get right, and no separate STT, LLM and TTS bills to reconcile. Step-by-step configuration, plus tools, voices and webhooks, is in the Twilio connection guide.

If Path A covers your use case, stop here and go build. The rest of this post is for the teams it doesn't.

Path B: build the bridge yourself with FastAPI

You want Path B when you need something the managed pipeline doesn't give you: a specific LLM you're contractually or architecturally committed to, a custom voice, turn-taking logic tuned to your domain, or an existing agent pipeline that already works and just needs a phone number bolted on.

The architecture is a WebSocket relay. Twilio opens a Media Streams connection to your server. Your server opens a second WebSocket to AssemblyAI. Audio flows caller → Twilio → you → AssemblyAI, transcripts come back, you call an LLM, you call a TTS provider, and synthesized mulaw goes back down the same Twilio socket.

Before you start

You need:

  • An AssemblyAI API key with access to Universal-3.6 Pro Realtime
  • A Twilio phone number with Voice enabled
  • An LLM API key — this example uses OpenAI, but you can route through the LLM Gateway to reach Anthropic, Google and others through one endpoint
  • A TTS provider that emits ulaw_8000
  • Python 3.11+
  • ngrok, to expose your local server during development
pip install fastapi uvicorn websockets python-dotenv openai elevenlabs twilio

Step 1: Return TwiML that opens a media stream

When a call arrives, Twilio hits your webhook and expects XML telling it what to do. <Connect><Stream> opens a bidirectional WebSocket to your server and keeps the call alive for its duration.

from fastapi import FastAPI, Request, Response

app = FastAPI()

@app.post("/twilio/voice")
async def voice(request: Request):
    host = request.headers["host"]
    twiml = f"""<?xml version="1.0" encoding="UTF-8"?>
<Response>
  <Connect>
    <Stream url="wss://{host}/media-stream" />
  </Connect>
</Response>"""
    return Response(content=twiml, media_type="application/xml")

Register this URL as the number's incoming voice webhook in the Twilio console, with the method set to POST.

Step 2: Bridge Twilio Media Streams to Universal-3.6 Pro Realtime

This is the core of the build. Twilio sends JSON frames with an event field: start when the stream opens (carrying the streamSid you need to send audio back), media for each 20ms chunk of base64-encoded mulaw, and stop when the call ends. Twilio also sends a dtmf event when the caller presses a key — more on that below.

Note the query parameters on the AssemblyAI socket. Streaming uses the singular speech_model parameter, unlike the async API which takes a plural speech_models list.

import asyncio, base64, json, os
import websockets
from fastapi import WebSocket

AAI_URL = (
    "wss://streaming.assemblyai.com/v3/ws"
    "?speech_model=universal-3-6-pro"
    "&encoding=pcm_mulaw"
    "&sample_rate=8000"
    
    "&voice_focus=near-field"
)

@app.websocket("/media-stream")
async def media_stream(twilio_ws: WebSocket):
    await twilio_ws.accept()
    stream_sid = None

    async with websockets.connect(
        AAI_URL,
        additional_headers={"Authorization": os.environ["ASSEMBLYAI_API_KEY"]},
    ) as aai_ws:

        async def pump_to_assemblyai():
            nonlocal stream_sid
            async for raw in twilio_ws.iter_text():
                msg = json.loads(raw)
                event = msg.get("event")

                if event == "start":
                    stream_sid = msg["start"]["streamSid"]

                elif event == "media":
                    audio = base64.b64decode(msg["media"]["payload"])
                    await aai_ws.send(audio)

                elif event == "dtmf":
                    await handle_dtmf(msg["dtmf"]["digit"], twilio_ws, stream_sid)

                elif event == "stop":
                    await aai_ws.close()
                    break

        async def pump_from_assemblyai():
            async for raw in aai_ws:
                msg = json.loads(raw)
                if msg.get("type") == "Turn" and msg.get("end_of_turn"):
                    transcript = msg.get("transcript", "").strip()
                    if transcript:
                        await handle_turn(transcript, twilio_ws, stream_sid)

        await asyncio.gather(pump_to_assemblyai(), pump_from_assemblyai())

Three things worth calling out.

format_turns=true is a legacy parameter that has no effect on the U3 Pro family. Universal-3.6 Pro Realtime always returns formatted, punctuated, cased text, so there is nothing to set here — you get it for free.

voice_focus=near-field isolates the primary speaker and suppresses background noise. near-field is the correct setting for headsets and phone handsets; far-field is for rooms, kiosks and drive-thrus. On a call from a car or a busy office this is the single highest-leverage flag on the socket.

End-of-turn detection combines semantic context with voice activity rather than waiting on a silence timer. That is why you can act on end_of_turn directly instead of layering your own VAD heuristic on top.

Step 3: Add agent_context — the highest-value change in this post

If you make one change to a phone agent's STT config, make it this one. agent_context lets you pass the agent's own last question into the stream, so the model knows what kind of answer to expect. When the agent asks "what's your zip code?", a mumbled "double-oh two one three" resolves correctly instead of becoming words.

Across a benchmark of 20,000 voice agent audio files, agent_context cut word error rate by 10.2%. The breakdown is where it gets interesting for phone agents specifically: fabrications down 18.3%, hallucinations down 17.2%, place-name entities down 15.5%, short-utterance errors down 13.7%, and name entities down 9.4%.

Short-utterance errors are the ones that kill phone agents. "Yes," "nine," "B as in boy" — these are the turns with the least acoustic information and the highest cost when they're wrong.

Send it as a config update on the socket right before you play the agent's question:

async def ask(question: str, aai_ws, twilio_ws, stream_sid):
    await aai_ws.send(json.dumps({
        "type": "UpdateConfiguration",
        "agent_context": question,
    }))
    await speak(question, twilio_ws, stream_sid)

Context Carryover — a short rolling memory of the conversation — is on by default and works alongside this, so you are not responsible for feeding back the whole history yourself.

You can also pick a latency posture with the min_latency, balanced and max_accuracy modes rather than tuning low-level flags. balanced is the default and the right starting point for a phone agent. Move to min_latency for terse, transactional flows where turns are short, and to max_accuracy when a turn is carrying something expensive to get wrong, like a payment reference.

Step 4: Run the LLM loop with tool calling

On each finalized turn, append the transcript to the conversation, call the LLM with your tool definitions, dispatch any tool calls, and feed the results back for a final response.

from openai import AsyncOpenAI

client = AsyncOpenAI()

TOOLS = [{
    "type": "function",
    "function": {
        "name": "get_order_status",
        "description": "Look up the status of a customer order by ID.",
        "parameters": {
            "type": "object",
            "properties": {"order_id": {"type": "string"}},
            "required": ["order_id"],
        },
    },
}, {
    "type": "function",
    "function": {
        "name": "transfer_to_human",
        "description": "Transfer the caller to a human agent.",
        "parameters": {
            "type": "object",
            "properties": {"reason": {"type": "string"}},
            "required": ["reason"],
        },
    },
}]

async def run_llm(messages):
    resp = await client.chat.completions.create(
        model="gpt-4o", messages=messages, tools=TOOLS, stream=False,
    )
    msg = resp.choices[0].message
    if not msg.tool_calls:
        return msg.content

    messages.append(msg)
    for call in msg.tool_calls:
        result = await dispatch(call.function.name, json.loads(call.function.arguments))
        messages.append({
            "role": "tool",
            "tool_call_id": call.id,
            "content": json.dumps(result),
        })

    follow_up = await client.chat.completions.create(
        model="gpt-4o", messages=messages,
    )
    return follow_up.choices[0].message.content

Keep tool handlers fast. A slow lookup does not just add latency, it adds silent latency, and callers hang up on silence long before they hang up on a slow answer. If a tool can take more than a second, say something first: "let me pull that up."

Step 5: Stream TTS back to Twilio as mulaw

Request ulaw_8000 from your TTS provider so the audio is already in Twilio's format. Base64-encode each chunk as it arrives and send it as a media event on the same socket, tagged with the streamSid you captured in the start event.

from elevenlabs.client import AsyncElevenLabs

eleven = AsyncElevenLabs()

async def speak(text: str, twilio_ws, stream_sid: str):
    stream = eleven.text_to_speech.stream(
        voice_id=os.environ["ELEVEN_VOICE_ID"],
        text=text,
        model_id="eleven_turbo_v2_5",
        output_format="ulaw_8000",
    )
    async for chunk in stream:
        await twilio_ws.send_text(json.dumps({
            "event": "media",
            "streamSid": stream_sid,
            "media": {"payload": base64.b64encode(chunk).decode()},
        }))

Twilio plays frames as they arrive, so the caller starts hearing the reply before the sentence has finished synthesizing. Send chunks as you get them — do not buffer the full response.

For barge-in, send Twilio a clear event on the same socket when a new Turn starts while you are still speaking. That flushes Twilio's playback buffer so the agent stops talking the instant the caller does.

Step 6: Handle DTMF

Phone agents need keypad input. Callers will punch in account numbers, menu selections and PINs whether or not you designed for it, and speech recognition is the wrong tool for a 16-digit card number read aloud in a parking lot.

Twilio sends a dtmf event on the Media Streams socket with the digit and the track. Buffer digits, terminate on # or on a short idle timeout, and route the collected string into your conversation state:

dtmf_buffer = []

async def handle_dtmf(digit: str, twilio_ws, stream_sid):
    if digit == "#":
        collected, dtmf_buffer[:] = "".join(dtmf_buffer), []
        reply = await run_llm(build_messages(f"[caller entered: {collected}]"))
        await speak(reply, twilio_ws, stream_sid)
    elif digit == "*":
        dtmf_buffer.clear()
    else:
        dtmf_buffer.append(digit)

The rule of thumb: collect anything long, numeric and high-stakes over DTMF, and everything else by voice. On Path A, DTMF is handled for you.

Step 7: Run it and connect Twilio

uvicorn server:app --port 8000
ngrok http 8000

Take the ngrok HTTPS URL, set the number's voice webhook to https://<your-ngrok-host>/twilio/voice with method POST, and call it.

Test streaming accuracy on your own call audio

Drop a recording of a real customer call into the playground and see how Universal-3.6 Pro Realtime handles your order numbers, names and addresses before you write any code.

Try playground

Latency budget: where your milliseconds go

Phone calls are unforgiving about silence. On a call there is no spinner, no typing indicator, and no visual cue that anything is happening — a pause reads as a dropped call, and callers start talking over the agent.

Here is what we can actually put numbers on:

  • End-of-turn detection. Universal-3.6 Pro Realtime decides a turn has ended with turn detection that combines semantic context with voice activity, rather than a fixed silence threshold. This matters more than the raw number: a silence timer set aggressively cuts people off mid-thought, and set conservatively it makes every turn feel slow. Judging completeness avoids that tradeoff.
  • SIP telephony: under 500ms. That is the managed Path A leg, down from the multi-second range earlier bridging approaches lived in.
  • Voice Agent API end-to-end: ~1 second. Caller stops speaking to caller hears audio, covering STT, LLM and TTS.

On Path B, your end-to-end number is the sum of your own choices, and we are not going to invent component figures for a pipeline we don't control. What we can tell you is where the budget actually goes wrong:

Resampling. If any hop in your chain converts 8kHz mulaw to something else and back, you pay for it twice and gain nothing. Keep mulaw end to end.

‍Non-streaming LLM calls. Waiting for a complete response before starting TTS creates a dead zone the length of the whole generation. Stream tokens and start synthesizing on the first clause boundary.

Slow tools. A database lookup that takes 800ms is 800ms of dead air unless you cover it. Acknowledge first, look up second.

Serialized TTS. Synthesizing the entire reply before sending the first frame wastes everything you saved upstream. Send chunks as they come.

Get those four right and the pipeline stays inside the range callers tolerate. Get any one of them wrong and no amount of model speed will save you. There is more on pipeline design in our voice agent architecture guide.

What about the AssemblyAI Voice Agent API?

This is the natural next step for most teams who build the Path B bridge and then get tired of maintaining it.

The Voice Agent API collapses the entire chain in this tutorial — STT, LLM, TTS, turn detection and tool calling — behind a single WebSocket. It is built on the same Universal-3.6 Pro Realtime model you were streaming to in Step 2, so the speech layer does not change; what changes is that you stop writing orchestration.

One connection. The endpoint is wss://agents.assemblyai.com/v1/ws. Note the base host: it is agents.assemblyai.com, not api.assemblyai.com. This trips people up regularly, and a connection failure here usually turns out to be the wrong host rather than a bad key.

One bill. Flat $4.50/hr ($0.075/min) covering speech-to-text, the LLM and text-to-speech together. No separate provider invoices to reconcile, and one set of logs to debug against instead of three. Full breakdown on the pricing page.

Standard JSON, no SDK required. It is a plain JSON protocol over a WebSocket, so it works with whatever you already have — including Claude Code out of the box.

Session resumption. Reconnect within 30 seconds and the conversation context is preserved. On mobile calls, where a connection blip is a normal event rather than an exception, this is the difference between a dropped call and a hiccup.

Six languages — English, Spanish, French, German, Italian and Portuguese — with native code-switching across all six. (Streaming STT on its own covers 32 languages with mid-sentence code-switching; the supported languages post has the full picture for both.)

Tool calling, on your side or ours. Tools can run client-side as in Step 4, or be configured over HTTP and executed on AssemblyAI's side. Parameter hints improve both accuracy and turn detection when an entity is actively being spoken, and tool-call hold mode covers the slow-lookup problem for you. See the tool calling docs.

Live configuration. Update the system prompt, the tool set or settings mid-conversation without reconnecting — useful for handing a call between stages of a flow.

There is also an Agent Management API for storing agent configs, webhooks for session and telephone-call lifecycle events, and output volume control. Background is in the launch post, and voice agent solutions covers how teams are deploying these.

The honest framing: if you are building a structured, transactional phone agent — support lookups, scheduling, ordering, qualification, triage — Path A and the Voice Agent API will get you there faster and cost you less to run. Path B earns its keep when you have a real constraint that the managed pipeline doesn't satisfy.

Where to go next

Start building your Twilio voice agent

Sign up for an API key and get free credit to test both paths — the managed SIP route and the DIY bridge — against your own call audio.

Sign up free

Frequently asked questions

How do I build a voice agent with Twilio and AssemblyAI?

There are two approaches. The fast one is managed SIP telephony: point a Twilio SIP endpoint at the AssemblyAI Voice Agent API, deploy one of the starter repos, and route your number to it — no bridge code required, and the SIP leg runs under 500ms. The custom one is a Media Streams bridge: return TwiML with <Connect><Stream>, forward Twilio's 8kHz mulaw frames to wss://streaming.assemblyai.com/v3/ws with speech_model=universal-3-6-pro, encoding=pcm_mulaw and sample_rate=8000, pass finalized turns to an LLM with tool definitions, and stream ulaw_8000 audio back on the same socket.

Should I use Twilio SIP or Twilio Media Streams?

Use SIP when you want the Voice Agent API to handle the whole pipeline — it is fewer moving parts, it runs under 500ms on the telephony leg, DTMF is handled for you, and billing is a single flat rate. Use Media Streams when you are bridging to your own stack and need control over which LLM, which voice, and how turns are orchestrated. SIP is the shorter path for most structured, transactional phone agents; Media Streams is the right call when you have an existing pipeline you can't or won't replace.

Why use AssemblyAI for a Twilio voice agent?

Universal-3.6 Pro Realtime accepts Twilio's native 8kHz mulaw in both directions, so there is no resampling stage adding latency and no format conversion to debug. On AssemblyAI's English voice-agent benchmark it posts a 2.4% error rate on phone numbers against Deepgram Flux EN's 11.5%, and 10.9% on names against 29.0% — the two categories phone agents live and die on. Turn detection combines semantic context with voice activity rather than waiting out a silence timer.

Does the Voice Agent API work with Twilio?

Yes. SIP telephony is shipped and plugs into any SIP endpoint, Twilio's included, with latency under 500ms. It works with AssemblyAI-owned numbers or numbers you already own, and DTMF keypad input is supported. Connect over wss://agents.assemblyai.com/v1/ws — note the base host is agents.assemblyai.com, not api.assemblyai.com.

Can a Twilio voice agent handle DTMF keypad input?

Yes, on both paths. With the Voice Agent API over SIP, DTMF is supported natively. On a Media Streams bridge, Twilio sends a dtmf event on the same WebSocket carrying the digit and track — buffer the digits, terminate on #, and inject the collected string into your conversation state. Collect long numeric values like account numbers and card references over DTMF rather than speech; it is faster for the caller and it removes an entire class of recognition risk.

What latency should I expect from a Twilio voice agent?

On the managed path, expect under 500ms on the SIP telephony leg and roughly one second end-to-end from the caller finishing their sentence to hearing a reply. On a DIY bridge the number depends on the LLM and TTS you choose, but the four things that reliably wreck the budget are resampling audio between formats, waiting for a complete LLM response before starting TTS, slow tool calls that produce dead air, and buffering the full TTS output before sending the first frame.

How much does it cost to run a phone-based voice agent?

On the managed path it is a flat $4.50/hr ($0.075/min) for STT, LLM and TTS combined, plus Twilio's per-minute voice charges, which vary by country. On a DIY bridge you pay Streaming Speech-to-Text at $0.45/hr base plus your own LLM and TTS providers plus Twilio. The flat rate is usually cheaper once you account for the engineering time spent maintaining a bridge; current numbers are on the pricing page.

Can a Twilio voice agent handle multiple simultaneous calls?

Yes. AssemblyAI streaming sessions are billed per session-hour and streaming rate limits scale with your usage, so check your account's concurrency headroom rather than assuming there is no quota. Twilio applies concurrency limits per account, which you can raise. On a DIY bridge, each call is one FastAPI WebSocket handler and one upstream AssemblyAI socket, so a single uvicorn process handles hundreds of concurrent calls comfortably — run several behind a load balancer and scale horizontally from there.

‍

Title goes here

Lorem ipsum dolor sit amet, consectetur adipiscing elit, sed do eiusmod tempor incididunt ut labore et dolore magna aliqua. Ut enim ad minim veniam, quis nostrud exercitation ullamco laboris nisi ut aliquip ex ea commodo consequat. Duis aute irure dolor in reprehenderit in voluptate velit esse cillum dolore eu fugiat nulla pariatur.

Button Text
AI voice agents
Voice Agent API