New Universal-3.5 Pro is here. Learn more: Async Realtime
Voice agents · Customer support

Voice agents for customer support

Customer support is the highest-volume, highest-stakes use case for voice agents — and the one where accuracy failures are most expensive. Every misheard account number is a repeat call, every botched escalation a lost customer. Here's how to ship production-grade support experiences, from first-call resolution to real-time agent assist.

The production gap

Confidence vs. reality

82% of builders feel confident handling interruptions

55% of end users are frustrated by how agents handle them

The gap between shipped and actually works. Source: The Voice Agent Report.

The numbers that matter

Accuracy over cost

76%

Of organizations prioritize accuracy over cost when choosing voice AI infrastructure — because errors compound into repeat calls and churn.

Source: The Voice Agent Report

Fewer tickets

90%

Reduction in customer complaints and support tickets after Siro built on AssemblyAI.

Source: Siro case study

Phone-number error

3.55%

Universal-3.5 Pro Realtime's phone-number entity error rate on the Pipecat benchmark — best of the realtime models tested.

Source: Daily's Pipecat STT benchmark

End-of-turn

~300ms

Turn detection from tonality, pacing, and rhythm — so the agent doesn't cut callers off mid-sentence.

Source: Universal-3.5 Pro Realtime

Customer story · Report 2

“The transcription accuracy, reliability, and speed of AssemblyAI's API have greatly enhanced our operations, reinforcing our trust in their technology and solidifying our partnership.”

— Raj Shankar, SVP Product, Calabrio

The production gap

The production gap in customer support voice agents

Support agents in production face what demos never show: callers with accents and background noise, alphanumeric soup (order numbers, confirmation codes, emails), multi-turn conversations with context dependencies, and the need to escalate gracefully to a human. Accuracy failures don't stay small — they compound into repeat calls, CSAT drops, and churn.

Hears the caller, not the room

voice_focus isolates the primary speaker and suppresses background speech that would otherwise be transcribed as phantom words and fire false interruptions.

Universal-3.5 Pro Realtime

Async-grade speaker labels, live

SpeakerRevision labels speakers during the call, then corrects them with a single ~500ms revision at stream end (up to 10 speakers) — ready for supervisor dashboards.

Universal-3.5 Pro Realtime

Coaching in the moment

Live transcription feeds sentiment flags, escalation detection when calls go off-script, and post-call analytics via the Speech Understanding API.

Speech Understanding

Turn-based agents

IVR and call routing

For flows where you do your own turn detection and just need a finished transcript per utterance — IVR menus, call routing, one-shot voice commands — the Sync API returns a completed transcript in a single HTTP request. No WebSocket to manage.

Sync API flow

Low latency
One HTTP request
Transcript → finish the flow

Runs on Universal-3.5 Pro.

Two paths

Two paths to production support

Both run on the same Universal-3.5 Pro Realtime foundation — start fresh with the managed API, or slot the STT layer into your existing contact center stack.

Recommended

Voice Agent API

The fastest path to production support agents

$4.50 /hr

STT + LLM + TTS included — free tier available

  • One WebSocket, flat rate — no per-token surcharges
  • Built-in tool calling for account and order lookups
  • Session resumption — 30s reconnect, context preserved
  • Speech-aware end-of-turn (~300ms)
  • Best for: teams building new support agents
Get started

Bring your own stack

Universal-3.5 Pro Realtime STT

The BYO layer for existing contact centers

$0.45 /hr

Transcription only — unlimited concurrency

  • Slots into LiveKit, Pipecat, Vapi, Twilio SIP
  • voice_focus suppresses background speech and false interruptions
  • Dynamic prompting across call stages (verify → troubleshoot → escalate)
  • Native code-switching across supported languages
  • Best for: teams with existing contact center infra
Learn more

From pilot to scale

What production actually requires

Sophisticated teams want modular infrastructure they can plug into their own pipeline. Both lanes serve this — but production has table stakes.

Production readiness

Table stakes

  • Entity accuracy Pass
  • Unlimited concurrency Pass
  • Compliance (SOC 2 / BAA) Pass
  • Escalation to human Pass
  • Session resumption Pass
  • Entity accuracy on the numbers and names that matter. 3.55% phone-number error on the Pipecat benchmark — best of the realtime models tested.

  • Unlimited concurrency without rate-limit surprises. Scale to peak call volume without throttling.

  • Enterprise compliance. SOC 2 Type 2, ISO 27001:2022, PCI DSS v4.0, and a BAA for healthcare-adjacent support.

  • Graceful escalation to a human. Tool calling plus escalation detection, with full context handed off.

  • voice_focus for noisy call-center floors. Background speech doesn't become phantom words.

  • Session resumption for dropped mobile calls. Reconnect within 30 seconds, context preserved.

  • Proven in production — CallRail added voice agents and Siro cut complaints and support tickets by 90%.

Bring voice agents to your contact center

Get your free API key and start building — or talk to us about an enterprise contact center deployment.

Frequently asked questions

What is the best speech-to-text API for contact centers?

AssemblyAI is built for contact-center audio: Universal-3.5 Pro Realtime posts the best entity error rates of the realtime models on Daily's Pipecat benchmark (3.55% on phone numbers), isolates the caller from background floor noise with voice_focus, and runs with unlimited concurrency so peak call volume never hits a rate limit. Use it as a standalone STT layer, or through the fully managed Voice Agent API.

How does real-time agent assist work with a voice AI API?

Stream the live call to Universal-3.5 Pro Realtime and feed the transcript into your coaching layer: sentiment flags for frustrated callers, escalation detection when a conversation goes off-script, and live speaker labels corrected to async-grade accuracy with a single ~500ms SpeakerRevision at stream end. Post-call, the Speech Understanding API adds entity detection, topic classification, and summarization.

Can a voice agent transfer a call to a human?

Yes. The Voice Agent API supports tool calling and escalation detection, so the agent can hand off to a human the moment a conversation needs it — with the full transcript and context passed along so the caller never has to repeat themselves.

What is the difference between a voice agent and an IVR?

An IVR follows a fixed menu tree ('press 1 for billing'); a voice agent understands natural speech and can look up an account, answer a question, or resolve an issue in conversation. For simple turn-based flows like IVR menus or call routing, the Sync API returns a finished transcript per utterance in a single HTTP request — no WebSocket to manage.

How do you keep a support call from restarting when the connection drops?

The Voice Agent API's session resumption reconnects within 30 seconds with the full conversation context preserved — so a dropped mobile call or a flaky network doesn't restart the interaction.

How much does it cost to run a support voice agent?

The Voice Agent API is a flat $4.50/hr with STT, LLM, and TTS included and no per-token surcharges. If you already run your own orchestrator, Universal-3.5 Pro Realtime is $0.45/hr for transcription only, with add-ons (diarization, prompting, voice isolation) billed only as used.