How to build a multilingual voice agent with the Voice Agent API
A hands-on build guide: connect over one WebSocket, steer languages, pass conversation context, add tools, and test a multilingual voice agent end to end.



This is a build guide. By the end you'll have a voice agent that answers in more than one language, keeps up when a caller switches mid-sentence, calls your backend when it needs a real answer, and survives a dropped connection without making the caller start over.
The whole thing runs over one WebSocket. No SDK, no orchestration framework, no three-vendor pipeline to keep in sync. If your language can open a socket and parse JSON, you can build this.
If you're still deciding whether to build one — the architecture, the benchmarks, how to evaluate a vendor — the concept guide covers that ground. This post assumes you've decided.
What you'll build
A voice agent for a delivery company's support line. It greets callers, understands them in English, Spanish, or French without asking which they'd prefer, looks up an order when asked, and handles a caller who says "mi pedido es one-four-seven-two" without falling over at the language boundary.
Six steps:
- Open a session and configure the agent
- Steer the languages
- Tune transcription for the room and the vocabulary
- Give the agent a tool
- Handle reconnects
- Test across languages
Before you start
You need three things:
- An API key. Free to create, no card required.
- Python 3.9+ and the websockets package. pip install websockets. That's the only dependency in this guide.
- A source of 24 kHz mono PCM16 audio. The Voice Agent API's default audio/pcm encoding is 24,000 Hz, 16-bit signed little-endian, mono — for both the audio you send and the audio you get back. A microphone in the browser, a Twilio media stream, or a WAV file you read off disk all work, as long as you resample to 24 kHz first. Start with the file — it makes everything reproducible. This is the single most common setup mistake: 16 kHz is the reflex sample rate for streaming APIs, and feeding 16 kHz audio to this endpoint produces sped-up, chipmunked audio and garbage transcripts with no error to tell you why.
Budget-wise, the Voice Agent API is a flat $4.50/hr covering transcription, the language model turn, and speech synthesis together. No per-component metering to reconcile. See pricing for how that lands against composing it yourself.
How the Voice Agent API is wired
One WebSocket carries everything: your audio up, the agent's audio down, and every transcript, tool call, and lifecycle event as JSON in between. There's no separate transcription connection to manage and no LLM endpoint to call yourself.
Connect to wss://agents.assemblyai.com/v1/ws with your API key as a Bearer token on the upgrade request. Then the flow is:
| Direction | Event | What it's for |
|---|---|---|
| You send | session.update |
Configure the agent. First message, and again any time you want to change something mid-call. |
| You receive | session.ready |
Session is live. Save the session_id. Don't send audio before this. |
| You send | input.audio |
Base64-encoded 24 kHz mono PCM16, streamed in chunks, in the audio field. |
| You receive | transcript.user |
The caller's finalized turn. Deltas arrive first as transcript.user.delta. |
| You receive | reply.audio |
The agent's speech, base64 24 kHz PCM16, in the data field — not audio. Decode and play immediately. |
| You receive | tool.call |
The agent wants your backend. Reply with tool.result. |
| You receive | session.ended |
The session is finished — including after the session.end you send yourself. Always handle it: it's your cue to flush logs, stop the audio pump, and release the connection. |
That's the whole protocol surface you need for this build. The API reference has the exhaustive list.
Everything in this guide runs on a free account. Create one, grab your API key, and start the first session.
Step 1: Open the session
Connect, send one session.update, wait for session.ready, then start streaming audio. Here's the skeleton everything else plugs into.
import asyncio
import base64
import json
import os
import websockets
URL = "wss://agents.assemblyai.com/v1/ws"
API_KEY = os.environ["ASSEMBLYAI_API_KEY"]
# output.voice is immutable for the life of a session, so pick the voice that
# matches the language this session will reply in. One session per output
# language — see the note below.
VOICES = {
"en": ("alba", "English", "Hi, thanks for calling. How can I help?"),
"es": ("lola", "Spanish", "Hola, gracias por llamar. ¿En qué puedo ayudarle?"),
"fr": ("estelle", "French", "Bonjour, merci de votre appel. Comment puis-je vous aider ?"),
}
REPLY_LANGUAGE = "es"
voice, language_name, greeting = VOICES[REPLY_LANGUAGE]
SESSION_CONFIG = {
"type": "session.update",
"session": {
"system_prompt": (
"You are a support agent for a delivery company. "
f"Always reply in {language_name}, whatever language the caller uses. "
"Keep every reply to one or two short sentences."
),
"greeting": greeting,
"output": {"voice": voice},
},
}
async def main():
headers = {"Authorization": f"Bearer {API_KEY}"}
async with websockets.connect(URL, additional_headers=headers) as ws:
await ws.send(json.dumps(SESSION_CONFIG))
async for raw in ws:
event = json.loads(raw)
kind = event["type"]
if kind == "session.ready":
print("session:", event["session_id"])
# start pumping microphone or file audio here
elif kind == "transcript.user":
print("caller:", event["text"])
elif kind == "transcript.agent":
print("agent:", event["text"])
elif kind == "reply.audio":
# the payload lives in "data", not "audio" — the reverse of
# what you send on input.audio. 24 kHz mono PCM16.
play(base64.b64decode(event["data"]))
elif kind == "session.ended":
print("session ended")
elif kind == "session.error":
print("error:", event["code"], event["message"])
asyncio.run(main())Audio goes up as base64-encoded PCM16 mono, one message per chunk:
import base64
async def send_chunk(ws, pcm_bytes):
# pcm_bytes must already be 24 kHz, 16-bit signed little-endian, mono.
# This function base64-encodes whatever you hand it — it cannot detect a
# wrong sample rate, and 16 kHz audio will play back chipmunked.
await ws.send(json.dumps({
"type": "input.audio",
"audio": base64.b64encode(pcm_bytes).decode(),
}))Two rules that bite people. First, stream at real time or slower. Firing a whole WAV file at the socket as fast as it reads returns an audio_rate_violation. Second, watch the field names in each direction: you send audio in audio, and you receive it in data. That asymmetry is a trap, and a b64decode of the wrong key fails silently as a KeyError in the middle of a live call.
One session per output language
output.voice is immutable for the life of a session. You can change the input languages mid-call, the transcription prompt, the keyterms, and the system prompt — but not the voice. That single constraint shapes the architecture of any genuinely multilingual line.
So don't set an American English voice and then ask the system prompt to "reply in the caller's language." Transcribing Spanish while replying in an English voice is jarring, and the docs warn against exactly that. Pick the output language up front — from the number dialed, the caller's account record, or a brief language-selection turn — and open the session with the matching voice: alba for English, lola for Spanish, estelle for French. If the caller turns out to want a different reply language, you end that session and open a new one with the right voice, carrying your own conversation summary across.
Input is a different story, and this is the part that stays genuinely multilingual within one session: the model handles a caller code-switching mid-sentence no matter which voice you chose. It's only the reply that's pinned.
Step 2: Steer the languages
session.input.language_codes declares which languages to expect. Omit it and the model detects automatically. For a multilingual agent, declaring a short list is usually better than leaving it wide open — you narrow the search space without committing to one language.
"input": {
"language_codes": ["en", "es", "fr"],
},Add that to the session object from step 1. Now the agent expects English, Spanish, and French, and handles a caller moving between them mid-sentence.
You can change it mid-call. Send another session.update and the new list applies on the next speech-to-text connect — not the current turn, which is worth knowing if you're switching in response to something the caller just said.
await ws.send(json.dumps({
"type": "session.update",
"session": {"input": {"language_codes": ["es"]}},
}))Be precise about which layer you're configuring, and in which direction. The Voice Agent API understands 18 languages on input with native code-switching — English, Spanish, German, French, Portuguese, Italian, Turkish, Dutch, Swedish, Norwegian, Danish, Finnish, Hindi, Vietnamese, Arabic, Hebrew, Japanese, and Chinese — and speaks 6 of them back today with native-accent voices: English, Italian, Spanish, German, Portuguese, and French. Native-accent voices for the remaining twelve are on the roadmap. The underlying speech model, Universal-3.6 Pro Realtime, transcribes those 18 plus 14 more (32 languages in all) with mid-sentence code-switching when you use it directly.
So language_codes can name any of the 18 — recognition isn't your constraint. What constrains you is which language the agent can answer in, and, because of the immutability rule above, how many sessions you need to cover those answers. The three in this build, en/es/fr, sit inside both lists. If you need to reply in a language outside the six, use the streaming model directly under your own orchestration and voice provider, which the last section covers.
Step 3: Tune transcription for the room and the vocabulary
Three settings do most of the work here, and all three live under input.
"input": {
"language_codes": ["en", "es", "fr"],
"transcription_mode": "balanced",
"transcription_prompt": (
"Customer support calls about parcel delivery. Callers give "
"order numbers as digits and may switch between English, "
"Spanish, and French mid-sentence."
),
"keyterms": ["Parcelo", "Express Saver", "depot"],
"voice_focus": "far-field",
},transcription_mode trades speed against accuracy: min_latency, balanced (the default), or max_accuracy for noisy and far-field audio.
transcription_prompt is free-form context, up to 1,750 characters, and it's the most underused lever on this list. Telling the model that callers switch languages and read out digits measurably changes what it produces.
How much detail is worth writing? The streaming docs benchmark three prompt levels against sending nothing at all, across 20,000 real voice-agent calls. A bare domain label of two to five words ("Parcel delivery support call.") cuts word error rate about 5%. A scenario sentence of five to fifteen words — what the conversation is actually about — cuts WER about 10% and takes name and place entity errors down 16% and 21%. A detailed prompt of twenty to fifty words, naming the people, products, and identifiers in play, nearly halves name errors at −49%, with overall entity errors down 29% and place names down 44%. Scenario context is the sensible default, because it asks for nothing your application doesn't already know; detailed context is what you reach for when you have the caller's record in hand. Two results worth knowing before you worry about overdoing it: hallucinated words go down as you add context, not up, and turn detection is unaffected. The full table is in the streaming prompting documentation.
keyterms boosts exact strings. On the Voice Agent API, input.keyterms accepts up to 100 of them. Brand names, product tiers, internal jargon. Multilingual audio makes this more valuable, not less, because your product names are the tokens least likely to exist in the model's training data for any given language.
voice_focus isolates the primary speaker and suppresses everything behind them. Use near-field for headsets and phones, far-field for rooms and kiosks. Background chatter in a second language is a genuinely nasty failure mode, and this is what handles it.
Keyterms, mode, and prompt are all mutable mid-session. Voice focus is applied on the next speech-to-text connect.
Step 4: Give the agent a tool
Tools are declared as JSON Schema in session.tools. When the agent decides it needs one, you get a tool.call; you run it and send back a tool.result.
"tools": [
{
"type": "function",
"name": "lookup_order",
"description": "Look up a delivery order by its order number.",
"parameters": {
"type": "object",
"properties": {
"order_number": {
"type": "string",
"description": "The order number, digits only",
},
},
"required": ["order_number"],
},
}
],Handle the call in your event loop. Accumulate on tool.call and send results when reply.done arrives:
pending = []
if kind == "tool.call":
pending.append(event)
elif kind == "reply.done":
for call in pending:
result = lookup_order(call["arguments"]["order_number"])
await ws.send(json.dumps({
"type": "tool.result",
"call_id": call["call_id"],
"result": json.dumps(result),
}))
pending.clear()arguments arrives as a dict, ready to use. Write tool descriptions in English regardless of the languages you support — the model reasons over the schema, not the caller's language.
One thing worth knowing before you put this near a real backend: the Voice Agent API applies guardrails to tool calling that constrain arguments to what is actually inferable from the caller's turns and from prior tool results, rather than letting the model invent a plausible-looking order number. That's the difference between an agent that occasionally looks up the wrong order and one you can point at production. If you'd rather bring your own model for the reasoning step, the LLM Gateway handles that.
Step 5: Handle reconnects
If the socket drops, reconnect within 30 seconds and send session.resume with the saved session_id. The conversation context survives. The caller carries on rather than re-explaining themselves to something that just forgot them.
session_id = None
async def connect():
global session_id
headers = {"Authorization": f"Bearer {API_KEY}"}
async with websockets.connect(URL, additional_headers=headers) as ws:
if session_id:
await ws.send(json.dumps({
"type": "session.resume",
"session_id": session_id,
}))
else:
await ws.send(json.dumps(SESSION_CONFIG))
async for raw in ws:
event = json.loads(raw)
if event["type"] == "session.ready":
session_id = event["session_id"]
elif event["type"] == "session.error" and event["code"] in (
"session_not_found", "session_forbidden", "session_expired"
):
session_id = None # dead session, start freshAll three resume failures have to clear session_id, not just the first two. session_expired is the one people miss: if the session's TTL elapses while you're inside the 30-second grace window, that's the code you get back, and a handler that only matches session_not_found and session_forbidden will hold onto a dead id and re-send the same doomed session.resume on every reconnect thereafter.
When the call genuinely ends, send {"type": "session.end"}. You'll get a session.ended back to confirm. Just closing the socket leaves the session held open for the 30-second grace window, and that window is billable.
Step 6: Test across languages
Here's where multilingual agents get shipped broken, because the obvious test — one clean recording per language — passes easily and predicts nothing.
Test these five, in this order:
- Clean single-language calls. Your baseline. If these fail, nothing else matters.
- Code-switched calls. Record a bilingual speaker doing what they naturally do — an English product name inside a Spanish sentence, a French caller spelling a German surname.
- Short replies. One-word and two-word turns in every language: "sí," "that's right," "quatre." Short utterances are where agents fail and where nothing in your aggregate metrics will warn you.
- Entities under noise. Order numbers, addresses, and names read aloud over background audio. Flip voice_focus and transcription_mode and measure the difference rather than guessing.
- Turn-taking per language. Count how often the agent interrupts and how long it waits after the caller stops, separately for each language.
Log transcript.user against your reference text and score entities separately from words. A 6% word error rate that misses every order number is worse than an 8% one that gets them all.
Fireflies went through this evaluation for their own agent pipeline:
"We were searching for the best realtime ASR model for our voice agent pipeline in Fireflies. The new Universal 3.5 Pro speech model from Assembly is best so far in terms of accuracy, latency and language switching." — Foysal Osmany, Software Engineer at Fireflies
Talk to an agent in the playground, switch languages mid-sentence, and see how the turn detection responds.
Passing conversation context when you build your own stack
Inside the Voice Agent API, conversation context is handled for you — the agent's own replies and the caller's prior turns are already in the model's memory, so short answers resolve against the question that prompted them.
Assembling your own stack instead — our speech model plus your LLM and your voice provider — means passing the agent half yourself, and it's the single highest-value thing you can do for accuracy. The parameter is agent_context, sent to the streaming session after each agent reply:
await ws.send(json.dumps({
"type": "UpdateConfiguration",
"agent_context": "¿Cuál es su número de pedido?",
}))Now when the caller answers "uno cuatro siete dos," the model knows an order number is coming rather than guessing at four loose words. This is the same lever as the transcription prompt in step 3, applied per turn: the benchmark in that section is what to expect from it, and short replies are exactly where the gain shows up.
Set it at connect time to seed the opening greeting, then update after every reply, trimming long replies down to the substantive question — it's the question that does the work, not the pleasantries. The user side is carried forward automatically.
The keyterms limits change on this surface too. Streaming keyterms_prompt caps at 100 keyterms per session with each string 50 characters or shorter: go over 100 and the request errors, and individual terms longer than 50 characters are ignored. Both prompt and keyterms_prompt can be updated mid-stream through the same UpdateConfiguration message without reconnecting, and sending an empty array clears them. There's a Python streaming walkthrough if you're starting that path from scratch.
Where to go from here
You've got a working multilingual agent: language steering, contextual transcription, noise handling, tool calls, and reconnects. What's left is specific to your product.
Three directions worth taking next. Put it on a phone number, which is a SIP trunk rather than a media server. Move the config server-side as a stored agent so you're not shipping prompts in your client. And instrument the transcripts — you now have a per-language record of every turn, which is the raw material for finding out where your agent loses people.
The thing that surprises most teams at this point isn't the languages. It's how much of the perceived quality came from three lines of configuration — the transcription prompt, the keyterms list, and voice focus — rather than from the model choice they spent two weeks agonizing over.
For the architectural reasoning behind these choices and how to benchmark them against alternatives, the multi-language voice agent guide linked at the top is the companion to this one. Building with the Voice Agent API goes wider on non-multilingual features, the launch post covers the design decisions, and the changelog is where new session parameters show up first.
Everything in this guide runs on a free account at a flat $4.50/hr once you go live. Create a key and open your first session.
Frequently asked questions
How do I build a multilingual voice agent?
Open a WebSocket to the Voice Agent API, send a session.update with your system prompt and an input.language_codes list, then stream base64 24 kHz mono PCM16 audio and play back the reply.audio events you receive. Code-switching on input is handled by the speech model, so you don't route between per-language agents — but output.voice is fixed for the life of a session, so plan one session per language you reply in. The configuration that matters most for multilingual quality is the transcription prompt, the keyterms list, and voice focus.
What is AssemblyAI's Voice Agent API and how does it work?
It's a single WebSocket that handles speech-to-text, the language model turn, and speech synthesis together, at a flat $4.50/hr. You send JSON events up and receive JSON events down — no SDK required, and no separate providers to stitch together. End to end it responds in roughly 1 to 1.3 seconds, and you can reconfigure a live session without reconnecting.
How would I use streaming speech-to-text in my voice agent?
You have two options. Use the Voice Agent API and streaming transcription is included and managed for you, or connect to the streaming endpoint directly and drive your own LLM and voice synthesis. Take the second path when you need the agent to answer in a language the managed pipeline doesn't speak yet, or when you already run an orchestrator, and pass agent_context after each agent reply so short caller responses resolve correctly.
How do I build a voice agent with function calling?
Declare each function as JSON Schema in session.tools with a name, a description, and typed parameters. When the agent decides to call one you receive a tool.call event with the arguments already parsed into a dict, and you reply with a tool.result carrying the same call_id. Write tool names and descriptions in English even for a multilingual agent, since the model reasons over the schema rather than the caller's language.
Can I build a language learning voice agent that handles code-switching?
Yes, and it's one of the few cases where code-switching is the product rather than an edge case. List both the target language and the learner's native language in language_codes so the agent follows when a learner drops back mid-sentence. Use the transcription prompt to describe the exercise, which helps the model expect hesitant and non-native pronunciation.
What happens if the connection drops during a call?
Reconnect within 30 seconds and send session.resume with the session_id you saved from session.ready, and the conversation context is preserved. If the resume fails you'll get a session_not_found, session_forbidden, or session_expired error — handle all three by clearing the saved id and starting a fresh session. Send session.end when a call is genuinely over, since simply closing the socket keeps the billable grace window open; you'll receive a session.ended event once it's torn down.
Lorem ipsum dolor sit amet, consectetur adipiscing elit, sed do eiusmod tempor incididunt ut labore et dolore magna aliqua. Ut enim ad minim veniam, quis nostrud exercitation ullamco laboris nisi ut aliquip ex ea commodo consequat. Duis aute irure dolor in reprehenderit in voluptate velit esse cillum dolore eu fugiat nulla pariatur.




