Skip to main content
Pick any voice ID from the tables below and set it as voice when you create the agent:
Configuring inline over the WebSocket instead of on a stored agent? Set it on session.output.voice in a session.update before session.ready:
The voice is immutable once the session is established, so it can’t be changed mid-conversation — pick it before the agent connects.

Language support

The agent recognizes far more input languages than it speaks. For the authoritative, up-to-date matrix, see Supported languages. In short:
  • Input (recognized): 18 languages, with native code-switching.
  • Output (spoken): 6 officially supported languages — 🇺🇸🇬🇧 English, 🇮🇹 Italian, 🇪🇸 Spanish, 🇩🇪 German, 🇵🇹 Portuguese, and 🇫🇷 French — each backed by at least one voice below with a matching accent. More output languages are on the roadmap.

Choose a voice by accent

Each voice has a primary accent:
  • For English — the right pick for an English agent — choose from English voices. Most have a 🇺🇸 American accent; paul and vera have a 🇬🇧 British accent.
  • For a native accent in Italian, Spanish, German, Portuguese, or French, choose the matching language-specific voice.

English voices

These voices have an English accent and are the recommended pick for English agents.

Language-specific voices

These voices have a native accent in a specific non-English language and code-switch naturally between that language and English.

Deprecated voices

The voices below are deprecated and will be removed in a future release. Use one of the Voices or language-specific voices above for new agents, and migrate existing agents when convenient.