There are six things being sold to you right now under the heading “AI in customer service,” and they are not the same kind of thing at all. Some of them are a weekend of integration work against call recordings you already have. Some of them put a machine on a live phone line with a paying customer on the other end.
Most articles on this topic rank them by how transformative they sound. That’s a bad ordering, because it puts the riskiest work first.
Here’s a better one. Sort the six by when the AI touches the conversation. Three of them run on calls that already ended. Three of them run while the caller is still talking. That single distinction predicts almost everything you care about as a buyer: what it costs, how long it takes to prove, how accurate the speech recognition underneath has to be, and what happens when it gets something wrong.
This post walks all six in that order, says what each one needs to work, and ends with the actual arithmetic on what a voice agent costs per call. If you’re the person who has to build the thing rather than fund it, the companion piece on Voice AI for customer service is the one you want — it goes a level deeper on the API surface and skips the portfolio question entirely.
What AI in customer service actually means
AI in customer service is the use of speech recognition and language models to understand, assist with, or handle customer conversations across phone, chat, and email. In practice that covers two distinct jobs: understanding conversations after they happen, which is analytics, and participating in conversations as they happen, which is automation.
The confusion in the category comes from the fact that both get the same name. A quality monitoring system that scores yesterday’s calls and a voice agent that answers the phone are both “AI in customer service,” and they share a foundation — accurate transcription — but almost nothing else. Different budget, different risk, different team.
So let’s separate them.
The three use cases that run on calls you already recorded
These are the ones to fund first, and it isn’t close. Your call recordings already exist. Nobody is waiting on the output in real time, so a slow answer is a fine answer. And if the model gets something wrong, the failure lands in a dashboard, not in a customer’s ear.
Automated quality monitoring and compliance
Most contact centers review something like one to two percent of calls, chosen by whoever has time. Every conclusion drawn from that sample is a conclusion about the one percent. Transcribe everything and score everything, and the sampling problem disappears — you’re no longer inferring what’s happening on your floor, you’re reading it.
This is the highest-certainty project on the list. Calabrio, which builds workforce optimization software for contact centers, reports an 80% increase in customer satisfaction on the back of this kind of work. They’re one of a group of contact center platforms building on our models, alongside Concentrix, CloudCall, AmplifAI, and Global Telesourcing.
The economics are unusually friendly. Transcribing recorded audio with Universal-3.5 Pro runs $0.21/hr. Ten thousand hours of recorded calls a month — a mid-sized center — comes to about $2,100 to transcribe in full. That’s the whole input cost for moving from one percent coverage to a hundred.
Customer intent and sentiment analysis
The second use case is the one that changes what leadership argues about in meetings. Once every call is transcribed, you can ask what customers are actually calling about, in their words, at volume — not what your disposition codes say they called about, which is a different and much less useful dataset.
The usual discovery is that two or three intents nobody has a process for account for a surprising share of volume. That finding is worth more than the tooling that produced it.
Predictive analytics and journey orchestration
This is the most oversold item in the category, so here’s the honest framing: it’s real, and it’s third for a reason. Predicting churn risk or next-best-action from conversation data works, but it needs the first two use cases running first, because it needs their output as its input. Fund it when you have twelve months of structured conversation data and not before. Buying it first is how organizations end up with an expensive model trained on disposition codes.
All three of these run on recorded audio, which means you can price the whole experiment today. Run a few hundred of your own call recordings through it — you’ll learn more about whether your audio quality supports this than any vendor conversation will.
The three use cases that run while the customer is still talking
Now the bar moves. Everything below happens live, which means latency matters, mistakes are visible to the customer, and the accuracy requirement changes shape in a way that’s worth being precise about.
AI-powered self-service voice agents
A voice agent answers the phone, understands what the caller wants, does something about it in your systems, and either resolves the call or hands it to a human with context attached. This is the use case that gets the coverage, and it’s the one with the most ways to go wrong.
The thing that makes or breaks it isn’t conversation design. Modern language models are good at conversation. What breaks it is whether the agent heard the order number.
Here’s why that distinction matters more than it sounds. On the Pipecat open STT benchmark, which runs on real agent conversations rather than read-aloud scripts, Universal-3.5 Pro Realtime posts a 6.99% word error rate against 15.58% for Deepgram Flux, 9.76% for ElevenLabs Scribe v2, and 9.04% for Google Chirp3. Respectable spread. But look at the entity error rate on the same benchmark: 15.31% for Universal-3.5 Pro Realtime, against 50.50%, 39.70%, and 21.51% respectively.
Entities are names, account numbers, addresses, and phone numbers — the things a support call exists to capture. A model can look fine on average word error rate and still mangle half the order numbers, because entities have no surrounding language to guess from. A customer never experiences your average. They experience the one field the agent got wrong, and then they ask for a human.
That’s the number to ask every vendor for, and most of them publish the other one. We put the comparison set on our benchmarks page so you can check the claim rather than take it.
Real-time agent assistance
Instead of replacing the human, this listens alongside them and surfaces the answer, the next question, or the compliance reminder while the call is live. It’s the highest-return live use case for most teams and it attracts the least attention, because it doesn’t make a good demo video.
Siro, which runs on our streaming speech-to-text, reports a 90% reduction in customer complaints and support tickets alongside a 36% improvement in close rate. Note what that first number is: fewer tickets, from better live conversations. The AI didn’t handle the contact. It made the human handle it better, which removed the follow-up contact entirely.
If you have to pick one live use case to start with, this is usually it. The failure mode is a bad suggestion that a trained human ignores. Compare that to a voice agent’s failure mode.
Intelligent routing and workflow automation
Routing on what the caller actually said in their first sentence, rather than on which menu key they pressed, is the least glamorous item here and one of the most reliable. It’s also the natural stepping stone: you’re running live speech recognition on real calls and acting on it, but a routing mistake costs a transfer, not a refund against the wrong account.
How voice agents work in customer service
Since the voice agent is the piece everyone asks about, here’s what’s underneath it — without turning this into an engineering document.
Core components and technical architecture
A voice agent is five components in a loop, running about once per conversational turn.
| Component | What it does | Current state, September 2026 |
|---|---|---|
| Phone system integration | Gets audio off your telephony or CCaaS platform and into the stack, and puts the reply back | Your existing carrier or CCaaS layer. Nothing here needs replacing. |
| Speech-to-text processing | Turns the caller’s audio into text, live, while they’re still speaking | Universal-3.5 Pro Realtime, 18 languages with native code-switching, $0.45/hr on its own |
| Conversational AI logic (LLM Gateway) | Decides what to say and which of your systems to call | One OpenAI-compatible endpoint across 37 models, with automatic fallback to a backup model |
| Text-to-speech synthesis | Turns the reply into speech the caller hears | Included in the Voice Agent API’s flat rate |
| Business system APIs | Looks up the order, issues the credit, books the appointment | Your CRM, order management, and ticketing systems, unchanged |
What changes when five components become one API
Until recently, rows two, three, and four meant three vendors. Three contracts, three bills, three sets of logs, and three places for latency to hide. When a call went badly, finding out why meant lining up timestamps across three dashboards that didn’t agree.
The Voice Agent API collapses those three rows into one connection at $4.50/hr flat, billed per second of session duration. Speech recognition, the model’s reasoning turn, and voice generation are one line item. It’s built on Universal-3.5 Pro Realtime, so the entity accuracy numbers above are the numbers you get.
For a CX leader the practical consequences are three. Procurement is one vendor instead of three. The bill is one number you can divide by call volume, rather than three that scale differently. And when something goes wrong on a call, there’s one place to look, which shortens every incident review you’ll ever run.
Rows one and five don’t collapse, and shouldn’t. Your telephony and your CRM stay yours. The engineering-level view of what that connection looks like in practice is in Voice AI for customer service, which is the piece to hand your platform team.
Integration with existing customer service systems
The question underneath “how does this integrate” is usually “do I have to rip out my contact center platform.” No. The voice agent sits in front of your existing stack and calls into it. Your CRM stays the system of record, your ticketing stays your ticketing, and your routing rules stay where they are — the agent becomes one more path into them. If audio residency or deployment location is a hard constraint for you, settle that before you go far down any vendor’s path — it decides which vendors are even eligible.
The integration work that’s genuinely hard isn’t the connection. It’s deciding what the agent is allowed to do without a human, which is a policy question your team has to answer regardless of which vendor you pick. The call analytics page covers what sits alongside an agent on the analytics side. Where stacks tend to hit their limits in production is its own subject, covered in where voice agent stacks show their limits.
Implementation considerations and success metrics
Choosing the right use case to start
The sequencing falls out of the split. Start with quality monitoring, because the data already exists and the failure mode is a wrong number in a report. Add real-time agent assistance second, because it’s live but a human stays in the loop. Add a voice agent third, scoped to one narrow high-volume intent — order status, appointment rescheduling, balance check — and not to “support.”
One thing we won’t do here is quote you a deflection rate. You’ll see a lot of them in this category, and none of them are measurements of your traffic. The percentage of your contacts a voice agent can resolve depends entirely on what your contacts are, and it’s knowable in about two weeks by running the agent against one intent. That’s a faster answer than any benchmark, and it’s the real one.
Measuring ROI and success metrics
Here’s the arithmetic, using the published flat rate.
The Voice Agent API is $4.50/hr, billed per second. A five-minute call costs $0.375. A four-minute call costs $0.30. A thousand calls a day averaging four minutes costs $300 a day, which is roughly $9,000 a month. If you need to model peak rather than average, a hundred concurrent calls running for eight hours costs $3,600.
On the analytics side, the arithmetic is different and smaller. Transcribing recorded calls for quality monitoring runs $0.21/hr on Universal-3.5 Pro, so ten thousand hours a month is about $2,100. Live speech recognition on its own, for agent assist, is $0.45/hr. The pricing page has the full table.
Now the number we can’t give you, and that you should be suspicious of anyone giving you: your fully loaded cost per contact today. Not the agent’s wage — the wage plus benefits, plus supervision, plus QA, plus the recruiting and training cost of replacing them, divided by contacts handled. Most centers know this number to within a few dollars, and it’s the only denominator that makes the comparison honest. Put $0.30 next to it and the business case either writes itself or it doesn’t.
Two metrics to track that aren’t cost. Containment with satisfaction held flat — containment alone is easy to fake by making escalation hard. And entity capture accuracy on your own traffic, because that’s the number that predicts whether customers will trust the thing.
Before you model any of this, drop a real recording into the playground — accents, hold music, background noise and all — and look specifically at whether the account numbers and names came through. An afternoon of that tells you more than a quarter of vendor evaluation.
Final words
The six use cases split into the three that run on yesterday’s calls and the three that run on today’s, and that ordering is most of the strategy.
But there’s a second-order effect worth naming, because it’s the thing teams discover a year in and wish they’d known. The analytics work isn’t just safer to do first — it’s what makes the automation work when you get there. Every intent you’ve catalogued, every product name and SKU format and account-number pattern you’ve pulled out of a year of transcripts, is exactly what a voice agent needs to be primed with to hear those things correctly on a live call. Teams who do quality monitoring first don’t just de-risk the automation. They arrive at it holding the vocabulary that makes it accurate.
Which means the boring project and the exciting one were never separate projects. The first one was the training data for the second.
If you’re moving a live support queue onto voice agents, the constraints worth settling early are concurrency, deployment location, and data residency. Those resolve faster in a conversation than in a documentation crawl.
Frequently asked questions
What’s the difference between AI chatbots and voice agents for customer service?
A chatbot works in text, where the input is unambiguous — the customer typed it, so the words are exactly what they meant. A voice agent has to recover the words from audio first, over a phone line, possibly with a car engine running. That extra step is where voice agents succeed or fail, and it’s why speech recognition accuracy matters far more than the conversational model. Voice agents are also synchronous: the caller is waiting, so a two-second pause reads as a broken system, while the same delay in chat reads as normal. The two also tend to cover different contacts, because people who pick up the phone usually have something more urgent or more complicated than people who open a chat window.
Which AI customer service feature should I implement first at my company?
Automated quality monitoring, in almost every case. Your call recordings already exist, so there’s no dependency on another team. Nothing runs in real time, so latency isn’t a constraint. Mistakes land in a report rather than in front of a customer. And at $0.21 per recorded hour to transcribe, you can cover a hundred percent of calls for roughly what sampling one percent costs in analyst time. It also produces the intent data and vocabulary that every later project needs. Real-time agent assistance is the sensible second step, and a narrowly scoped voice agent — one high-volume intent, not “support” — is the third.
How accurate does speech recognition need to be for customer service voice agents?
Ask about entity accuracy, not overall accuracy, because they come apart. On the Pipecat open STT benchmark, run on real agent conversations, Universal-3.5 Pro Realtime posts a 6.99% word error rate and a 15.31% entity error rate, against 15.58% and 50.50% for Deepgram Flux. The gap between those two columns is the point: a model can be within a few points on average words and be three times worse on names and account numbers, because entities have no surrounding language to infer from. Since a support call mostly exists to capture entities, that’s the column that determines whether customers trust the agent. Test on your own recordings and count how many account numbers came through exactly right.
Can AI customer service handle angry or frustrated customers effectively?
It can recognize frustration reliably and it should mostly respond by getting out of the way. Sentiment analysis on a live call is accurate enough to trigger an escalation rule, and the right rule is usually to route to a human early rather than to attempt de-escalation. A frustrated customer is frustrated partly because the last thing didn’t work, and an automated system insisting on one more attempt is a predictable way to make it worse. The place AI genuinely helps with angry customers is real-time agent assistance: the human takes the call, and the system surfaces the account history, the previous contact, and the policy they’re allowed to apply, so the human doesn’t spend the first ninety seconds reading. Design for graceful handoff with full context, not for automated de-escalation.
What happens if the AI customer service system doesn’t understand what I’m asking?
In a well-built system, it asks a narrower question once, and if that doesn’t resolve it, it transfers you to a person with the transcript of what you already said attached. The failure people hate isn’t being misunderstood — it’s repeating themselves to a human who has none of the context. Two design choices prevent most of it. Confirm anything consequential back to the caller before acting on it, especially numbers. And cap retries at one or two before handing off, rather than letting the system loop. On the recognition side, priming the model with your product names, SKU formats, and account-number patterns resolves a large share of the misunderstandings before they happen, because most of them are entities rather than intents.
Do we have to replace our contact center platform to add AI customer service?
No. Every use case in this post sits alongside the platform you already run rather than replacing it. Quality monitoring and intent analysis work off call recordings your platform is already producing. Agent assist runs in parallel with the live call. A voice agent sits in front of your existing telephony and calls into the same CRM and ticketing systems your human agents use, which stay the system of record. Several contact center platforms — Calabrio, Concentrix, CloudCall, AmplifAI, and Global Telesourcing among them — build these capabilities on our models rather than asking customers to move. The integration question that actually takes time isn’t technical; it’s deciding what the agent is allowed to do without a human reviewing it.