Insights & Use Cases
September 30, 2026

Build an AI voice agent for customer support that can look up orders

Build a Python voice agent that handles tier-1 customer support — order status lookups, account verification by email, and human escalation with conversation context — using AssemblyAI's Voice Agent API and tool calling on a single WebSocket.

Kelsey Foster
, 
Growth
Reviewed by
No items found.
Abstract green cylinder illustration
Table of contents

You deploy a customer support voice agent to a phone line by pointing a Twilio number at AssemblyAI over a SIP trunk and binding an agent ID to that number. There is no media server to run, no audio bridge to maintain, and no webhook in the call path — Twilio hands the call to AssemblyAI directly, and latency on that path is under 500ms. This guide covers the SIP setup, the audio and accuracy settings that matter at 8 kHz, DTMF keypad entry, escalation to a human, and the failure modes that only show up once real callers are on the line.

What this guide assumes you've already built

This is the deployment half of the story. It assumes you already have a working support agent: tools defined for order lookup and account verification, a system prompt that keeps replies short, and a session that runs end to end in a browser.

If you don't, build that first — how to build an AI voice agent for customer support that can look up orders walks through the tool definitions, the system prompt, and the WebSocket connection. Come back here when your agent answers correctly in a browser tab and you need it to answer a phone number instead.

The gap between those two states is larger than most teams expect. A browser gives you 24 kHz audio, echo cancellation, a stable connection, and a user who is looking at a screen. A phone line gives you 8 kHz μ-law, carrier jitter, a caller who may be driving, a keypad, and an expectation that a human is reachable within about ninety seconds. Everything below is about closing that gap.

Two ways to put an agent on a phone line

There are two paths, and the right one depends on whether you need to own the media.

SIP: the default

Twilio passes inbound calls to AssemblyAI over a SIP trunk. Your infrastructure is not in the audio path at all:

Caller → your Twilio number → SIP trunk → AssemblyAI → your agent

You register a number, bind an agent_id to it, and calls route. Nothing to host, nothing to scale, nothing to keep warm. It works with a number AssemblyAI provides or with a Twilio number you already own, which matters if you have an established support line you cannot move.

The Twilio Media Streams bridge

The alternative is the older pattern: Twilio opens a WebSocket to a service you run, that service relays audio to the Voice Agent API, and you relay the agent's audio back. The Voice Agent API speaks audio/pcmu — G.711 μ-law at 8 kHz — natively, so there is no transcoding step, but you are now running and scaling a media relay.

Set the encoding on the agent so the phone network's format goes straight through:

{
  "input":  { "format": { "encoding": "audio/pcmu" } },
  "output": { "format": { "encoding": "audio/pcmu" } }
}

audio/pcma (G.711 A-law) is also supported for carriers outside North America. Both default to audio/pcm at 24 kHz, so this is a change you have to make deliberately — leaving the default in place forces a resample on every frame in both directions.

Which one to pick

Choose SIP when Choose a Media Streams bridge when
You want inbound calls answered with no infrastructure of your own You need to fork, record, or inspect the media stream yourself
You're using a Twilio number you already own, or one AssemblyAI provides You're bridging a PBX, contact center platform, or carrier that isn't Twilio
Call setup latency matters — the SIP path is under 500ms, down from the 2–20 seconds a bridge-based setup could take You need outbound dialing, conferencing, or call control Twilio exposes only through its own APIs
You want one agent config to serve browser, app, and phone by ID You already run a media layer and adding one more leg is cheap

Most support deployments should start on SIP. Teams building on top of an existing UCaaS or contact center stack — the category business communications providers like Nextiva operate in — more often need the bridge, because the call already lives inside their own media plane before it reaches the agent.

Put an agent on a phone number today

The Voice Agent API bundles speech-to-text, LLM, and text-to-speech through one WebSocket at a flat $4.50/hr, with SIP telephony included. Start with the free tier.

Sign up free

Connect a Twilio number over SIP

Five steps: three on Twilio, two on AssemblyAI. Save your details first.

AAI_API_KEY="your-assemblyai-api-key"
AGENT_ID="7ad24396-b822-4dca-871a-be9cc4781cf9"    # the agent that answers
NUMBER="+1..."                                     # your Twilio number, E.164
TRUNK_DOMAIN="acme-support.pstn.twilio.com"        # you invent this

TRUNK_DOMAIN doesn't exist yet — you're naming the trunk you're about to create. It must end in .pstn.twilio.com and be unique across all of Twilio, so include your company name.

Create the trunk, route its origination to AssemblyAI, and attach your number:

# 1. Create the trunk
TRUNK_SID=$(twilio api:trunking:v1:trunks:create \
  --friendly-name "Support voice agent" \
  --domain-name "$TRUNK_DOMAIN" -o json | jq -r '.[0].sid')

# 2. Route incoming calls to AssemblyAI
twilio api:trunking:v1:trunks:origination-urls:create \
  --trunk-sid "$TRUNK_SID" --friendly-name "AssemblyAI SIP" \
  --sip-url "sip:sip.assemblyai.com" --priority 1 --weight 1 --enabled

# 3. Attach your number to the trunk
NUMBER_SID=$(twilio api:core:incoming-phone-numbers:list \
  --phone-number "$NUMBER" -o json | jq -r '.[0].sid')
twilio api:trunking:v1:trunks:phone-numbers:create \
  --trunk-sid "$TRUNK_SID" --phone-number-sid "$NUMBER_SID"

Then register the number with AssemblyAI and bind your agent to it:

AAI="https://agents.us.assemblyai.com/v1"

# 4. Register the number
curl -fsS -X POST "$AAI/phone-numbers/import" \
  -H "Authorization: Bearer $AAI_API_KEY" -H "Content-Type: application/json" \
  -H "Idempotency-Key: $(uuidgen)" \
  -d "{\"phone_number\":\"$NUMBER\",\"termination_uri\":\"$TRUNK_DOMAIN\"}"

# 5. Bind the agent
curl -fsS -X PUT "$AAI/phone-numbers/$NUMBER/agent" \
  -H "Authorization: Bearer $AAI_API_KEY" -H "Content-Type: application/json" \
  -d "{\"agent_id\":\"$AGENT_ID\"}"

Call the number. You should hear the greeting and be in a conversation.

Two things worth knowing before you debug anything. First, once the trunk controls the number, any Voice webhook still configured on the number itself no longer applies — a number that appears to do nothing when called is usually still pointed at an old webhook. Second, the binding is to the agent ID, not a snapshot of the agent. Edit the agent's prompt or tools and the next call picks up the change without touching Twilio. To swap agents entirely, re-run only the PUT.

The Python and JavaScript starter repos run all five steps as one idempotent command, so re-running is safe.

Tune the agent for phone audio

A phone line is a worse microphone than a laptop. 8 kHz μ-law throws away everything above about 3.4 kHz — which is exactly where the acoustic differences between S and F, M and N, and five and nine live. That is why order IDs and email addresses are the first thing to break when a browser agent moves to a phone.

Two settings and one benchmark matter here.

Isolate the caller

input.voice_focus strips background audio before it reaches transcription, which improves recognition and, more importantly, stops the agent from treating a car engine or a coworker as a turn:

{
  "input": {
    "voice_focus": "near-field",
    "voice_focus_threshold": 0.85
  }
}

near-field is the default and the right choice for handsets and headsets — the majority of support calls. Switch to far-field when callers are on speakerphone, in a car, or in a room mic. Match the model to how the caller is actually captured; guessing wrong costs more accuracy than leaving it alone.

Do not stack your own denoiser on top. The server already denoises, and a second layer adds artifacts that cost more than the noise did.

Know what accuracy you're actually getting

The Voice Agent API runs on Universal-3.6 Pro Realtime. Here is how it performs on AssemblyAI's English voice-agent benchmark, which scores 12,460 scripted voice-agent scenarios rather than clean read speech. Lower is better.

Metric Universal-3.6 Pro Realtime Deepgram Flux EN ElevenLabs Scribe v2 Deepgram Nova-3
Word error rate 5.19% 13.50% 7.78% 8.64%
Entity error rate 14.4% 30.1% 18.5% 26.1%
Names 10.9% 29.0% 14.8% 24.3%
Codes / IDs 10.0% 46.1% 12.0% 27.6%
Phone numbers 2.4% 11.5% 3.4% 4.5%

Two rows carry the weight for order lookup. Phone numbers at 2.4% is the digit-sequence case: order numbers, account IDs, ZIP codes, anything the caller reads out. Names at 10.9% is the harder one, and the reason to read values back before acting on them — a name is the least constrained field a caller will give you. On Pipecat's independent open STT benchmark of agent conversations, the same model posts a 0.96% pooled semantic word error rate. See the full benchmarks for methodology.

Give the model the conversation

Short answers are ambiguous in isolation. "It's B as in boy, four two" only parses correctly if you know the agent just asked for an order prefix. The Voice Agent API feeds the agent's own turns back into the speech model for you, which is why it holds up on clipped replies.

If you're running your own orchestration on Streaming Speech-to-Text instead of the bundled agent, you have to pass that context yourself, with agent_context:

await ws.send(json.dumps({
    "type": "UpdateConfiguration",
    "agent_context": "What's the order number? It starts with two letters."
}))

Across a benchmark of 20,000 voice agent audio files, agent_context cut word error rate by 10.2%, with the largest gains exactly where phone calls hurt: fabrications down 18.3%, hallucinations down 17.2%, place-name entities down 15.5%, and short-utterance errors down 13.7%.

On the Voice Agent API side, the equivalent lever you control directly is input.transcription_prompt, which biases transcription toward the vocabulary you expect:

{
  "input": {
    "transcription_prompt": "Caller is reading an order ID: two letters, a hyphen, five digits.",
    "keyterms": ["Acme", "expedited", "RMA"]
  }
}

Keyterm prompting is included at no extra charge on the Voice Agent API.

Use the keypad when the microphone won't do

Some values should never go through speech recognition. Card numbers are the obvious case, but long account IDs and claim numbers have the same problem: a caller reads sixteen digits, one gets dropped, and they hear "I didn't catch that" three times in a row.

DTMF solves this. The agent asks the caller to enter the value on the keypad, the digits arrive as tones rather than transcript, and for PCI-sensitive input they never enter the transcript, the logs, or the model at all. The dtmf agent in both starter repos demonstrates the full pattern.

Use DTMF for card entry and any digit sequence longer than about ten characters. Use speech for everything else — callers resent keypad menus, and the whole point of a voice agent is not being one.

If you do collect a long digit sequence by voice, write the tool's parameter pattern to tolerate how people actually speak. Callers read digits in groups, so the value often reaches your tool with spaces in it:

{
  "order_number": {
    "type": "string",
    "description": "Order number as spoken. May contain spaces between digits.",
    "pattern": " *([A-Z] *){2}- *([0-9] *){5}",
    "examples": ["AB-12345", "A B - 1 2 3 4 5"]
  }
}

A pattern that only accepts the tidy form rejects the spaced one and traps the caller in a re-ask loop. This is the single most common cause of a support agent that "keeps saying it didn't catch that."

Kill the dead air during a slow lookup

This is the failure mode the original version of this post could only describe. It now has a fix.

An order lookup that hits your order management system takes two or three seconds. In a browser that reads as a pause. On a phone line, two seconds of silence reads as a dropped call, and callers hang up or start talking over the agent.

Set execution_mode per tool to control what the agent does while a tool is in flight:

  • interactive (the default) — the agent speaks a natural transition phrase while the tool runs. Use it for database lookups, REST calls, anything returning in under about five seconds. This covers most order lookups.
  • hold — the agent goes silent and does not respond to the caller until you send a result. Use it for transfers, escalations, and long-running operations.

For a lookup that might exceed a few seconds, interactive with a generous timeout is the right answer, not hold. Wrapping a slow query in hold "to be safe" produces exactly the dead air you were trying to avoid.

When you genuinely need hold — a transfer, a payment authorization — you can still speak mid-hold without ending it, using reply.create:

await ws.send(json.dumps({
    "type": "reply.create",
    "instructions": "Tell the customer you're still pulling up their order."
}))

The hold continues until you send the matching tool.result. Sending tool.result auto-fires the next reply, so don't send both.

‍

One detail that catches people: during a hold, the server stops emitting live transcript events. Nothing is dropped — transcripts flush when the hold ends — but live captioning will appear to freeze if you're rendering it.

Move tool execution server-side

Client-side tools mean your relay has to be running, reachable, and correct for every call. HTTP tools move that round trip to AssemblyAI's side: you give a URL and a parameter schema, AssemblyAI makes the request and feeds the result to the model, and your client does nothing.

{
  "name": "get_order_status",
  "description": "Look up an order by its ID. Call this whenever the caller asks where their order is. Do not call this for returns or refunds.",
  "parameters": {
    "type": "object",
    "properties": {
      "order_id": {
        "type": "string",
        "description": "The order ID.",
        "pattern": "[A-Z]{2}-\\d{5}",
        "examples": ["AB-12345"]
      }
    },
    "required": ["order_id"]
  },
  "http": { "url": "https://api.example.com/orders", "http_method": "GET" },
  "execution_mode": "interactive",
  "timeout_seconds": 10
}

For a phone deployment this removes an entire class of outage. The number is bound to a stored agent, the agent calls your API directly, and there is no process of yours in the call path that can crash mid-conversation. See the tool calling docs for the full schema.

Try it before you wire up a number

Talk to a voice agent in the playground, tune the prompt and tools, then bind the same agent ID to a phone number.

Try playground

Escalation and warm transfer

Every support agent needs an exit. The difference between a good handoff and a bad one is whether the caller has to repeat themselves.

Model the transfer as a hold tool, not an interactive one — a transfer takes fifteen to thirty seconds, and an agent making small talk through it sounds evasive:

{
  "type": "function",
  "name": "transfer_to_human",
  "description": "Transfer the call to a human agent. Call this when the caller asks for a person, when a refund exceeds policy, or after two failed lookups.",
  "parameters": {
    "type": "object",
    "properties": {
      "department": {
        "type": "string",
        "enum": ["billing", "returns", "support"]
      },
      "summary": {
        "type": "string",
        "description": "One sentence describing what the caller needs."
      }
    },
    "required": ["department", "summary"]
  },
  "execution_mode": "hold",
  "timeout_seconds": 60
}

The summary parameter is what makes it warm. Your handler writes it — along with the order ID the agent already looked up and the session ID — into the ticket or screen-pop your human agents work from, so the person picking up starts with context instead of "how can I help you today?"

Send a reply.create status update about ten seconds in. Silence during a transfer is the moment callers hang up.

Gate the escalation tool behind the lookup tool rather than exposing both from the first turn. A tool that isn't in the current list can't be called, which means the agent cannot transfer before it has actually checked the order — and tool selection accuracy drops noticeably once an agent has more than about ten tools available at once.

Harden it before you point real traffic at it

Guardrails on tool arguments

A support agent that can issue a refund is a system where a hallucinated argument is a financial event, not a bad transcript. Tool parameter guardrails constrain arguments so they can only be inferred from user turns and tool results, never from the model's own generations. That took hallucinated tool arguments from 2.5% of 3,000 calls to 0%. Tool-call accuracy on an internal dataset improved from 69% to 80% over the same period.

Pair that with an explicit anti-fabrication clause in the system prompt. Guardrails stop the bad call; the prompt stops the bad claim:

NEVER state an order status, delivery date, refund amount, or confirmation
number unless that exact value came from a tool result in this conversation.
If you have not seen a tool result, you do not have these values. Do not
estimate. Do not say "around" a number.

Session resumption and the 30-second window

Mobile calls drop. If the client disconnects without sending session.end, the session stays open for 30 seconds so you can reconnect with session.resume and continue with context preserved — the caller does not re-explain anything.

The flip side is a billing gotcha worth knowing before your first invoice: that grace window is billable. Send session.end explicitly on any intentional disconnect — the caller hung up, someone hit "end call," the page unloaded — and skip it only when the disconnect was unintentional and you actually want the resume option.

ws.send(JSON.stringify({ type: "session.end" }));

Handle the three resume failures — session_not_found (the window expired), session_forbidden (wrong account), and session_expired — by starting a fresh session rather than retrying.

Concurrency

Your account has a concurrent-session limit on the Voice Agent API. Confirm your headroom before a product launch or a holiday spike, and handle concurrency_exceeded with backoff and a retry rather than dropping the call.

‍

What you do need to size is everything downstream. Your order management API is the constraint, not the agent. Set timeout_seconds per tool to something your backend can actually meet under load, and make sure a slow backend degrades into "let me get someone who can help" rather than a thirty-second silence.

What to log

Persist the session_id from session.ready for every call, not just the ones that go wrong. It is what support needs to locate a specific call. Alongside it, log the WebSocket close code and reason on disconnect, the session start timestamp, and whether you connected to the US or EU endpoint.

Webhooks give you the rest without polling. Subscribe a URL to session and call lifecycle events and AssemblyAI POSTs a signed payload when each one fires:

curl -X POST https://agents.assemblyai.com/v1/webhook-subscriptions \
  -H "Authorization: $ASSEMBLYAI_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "url": "https://your-app.com/webhooks/assemblyai",
    "events": ["call.connected", "call.ended", "call.failed", "session.completed"],
    "secret": "a-long-random-signing-secret-at-least-32-chars",
    "agent_id": "7ad24396-b822-4dca-871a-be9cc4781cf9"
  }'

call.ended carries the from and to numbers, the direction, the session ID, and links to the recording and transcript — enough to drive post-call QA, containment-rate reporting, and CRM writeback. call.failed is the one to alert on: it is how you find out a trunk misconfiguration is silently dropping calls.

Verify X-AAI-Signature (HMAC-SHA256 over the raw body) before trusting any delivery, and dedupe on event_id, because failed deliveries are retried with backoff.

Every call is also stored as a session with downloadable artifacts — the audio recording, and a timeline pairing each user_transcript with the agent_text that followed it. That timeline is the artifact to run evaluations against.

Output volume

output.volume takes a value from 0 to 100 and is adjustable mid-session. It sounds trivial until you A/B it on a real phone line: TTS tuned for laptop speakers is often too quiet through a handset earpiece, and callers read a quiet agent as a bad connection.

Where to go from here

The deployment is the easy half now. A SIP trunk, a bound agent ID, and the right audio encoding get you a working phone number in an afternoon. What takes longer is the operational work: reading back values before acting on them, gating destructive tools behind read tools, watching call.failed, and running evaluations against session timelines instead of guessing.

‍

The Voice Agent API bundles speech-to-text, LLM, and text-to-speech through one WebSocket at a flat $4.50/hr, which means one bill and one set of logs instead of three vendors to correlate when a call goes wrong. End-to-end latency runs about one second, and end-of-turn detection weighs whether the caller's sentence sounds complete rather than waiting out a silence timer — on the Sesame TurnBench evaluation that produces 55% fewer false interruptions than Streaming STT alone, 3.1 wrong interruptions per conversation versus 6.8.

It supports six languages — English, Spanish, French, German, Italian, and Portuguese — with native code-switching across all six, which covers most North American and Western European support lines but is worth checking against your actual call mix before you commit. See supported languages for the current list, and voice agent architecture for how the pieces fit together.

Planning a production support line?

Talk through call volume, escalation paths, and deployment options — including EU data residency and self-hosted — with someone who has done it.

Talk to AI expert

Frequently asked questions

How do I connect a voice agent to a phone number over SIP?

Create a SIP trunk in Twilio, point its origination URL at sip:sip.assemblyai.com, and attach your phone number to the trunk. Then register the number with AssemblyAI via POST /v1/phone-numbers/import and bind an agent to it with PUT /v1/phone-numbers/{number}/agent. There is no media server, audio bridge, or webhook in the call path — Twilio hands the call directly to AssemblyAI, and the number stays bound to the agent ID, so editing the agent takes effect on the next call without touching Twilio.

Does the AssemblyAI Voice Agent API support DTMF keypad input?

Yes. DTMF is supported, so callers can enter order numbers, account IDs, and card numbers on the keypad instead of speaking them. For PCI-sensitive input this is the recommended path, because keypad digits never reach the transcript, the logs, or the model. Use DTMF for card entry and long digit sequences, and speech for everything else.

What audio format should I use for a voice agent on a phone line?

Use audio/pcmu — G.711 μ-law at 8 kHz — which matches the phone network and avoids a resample on every frame. audio/pcma (G.711 A-law) is supported for carriers that use it. Both input and output default to audio/pcm at 24 kHz, so telephony deployments have to set the encoding explicitly. On the SIP path, Twilio and AssemblyAI negotiate this for you.

How do I stop a voice agent from going silent during a slow order lookup?

Set the tool's execution_mode to interactive, which is the default. The agent speaks a natural transition phrase while the tool runs, so a two-to-three-second order lookup doesn't read as a dropped call. For genuinely long operations that need hold mode, send a reply.create event with instructions to have the agent give a status update mid-hold without ending it.

How do I transfer a phone call from a voice agent to a human without dead air?

Define the transfer as a tool with execution_mode: "hold" and a timeout_seconds of about 60, and include a summary parameter the agent fills in. Your handler writes that summary — plus the order ID and session ID — into the ticket the human agent sees, so the caller doesn't repeat themselves. Send a reply.create status update roughly ten seconds into the transfer, since silence during a handoff is when callers hang up.

How many concurrent phone calls can an AssemblyAI voice agent handle?

Your account has a concurrent-session limit on the Voice Agent API, so confirm your headroom ahead of a launch or a seasonal spike. The practical limit is your own backend: set timeout_seconds per tool to what your order management API can meet under load, and make sure a slow backend degrades into an escalation rather than a long silence.

What should I log for a production voice agent deployment?

Persist the session_id from session.ready on every call, plus the WebSocket close code and reason, the session start timestamp, and which region you connected to. Subscribe to webhooks for call.connected, call.ended, call.failed, and session.completed to drive post-call QA and CRM writeback without polling, and alert on call.failed — that's how you detect a trunk misconfiguration dropping calls. Each session also stores a downloadable timeline pairing every user transcript with the agent reply that followed it.

What happens if a call drops mid-conversation?

If the client disconnects without sending session.end, the session stays open for 30 seconds and you can reconnect with session.resume to continue with full context preserved. That grace window is billable, so send session.end explicitly whenever a disconnect is intentional. Handle session_not_found, session_forbidden, and session_expired by starting a fresh session rather than retrying the resume.

Title goes here

Lorem ipsum dolor sit amet, consectetur adipiscing elit, sed do eiusmod tempor incididunt ut labore et dolore magna aliqua. Ut enim ad minim veniam, quis nostrud exercitation ullamco laboris nisi ut aliquip ex ea commodo consequat. Duis aute irure dolor in reprehenderit in voluptate velit esse cillum dolore eu fugiat nulla pariatur.

Button Text
AI voice agents
Voice Agent API