New Universal-3.6 Pro Realtime is now available Learn more

The ongoing engineering cost of a multi-vendor voice agent stack

What nobody tells you during the evaluation phase: the real operational surface area of running STT, LLM, TTS, and orchestration as a single product.

The ongoing engineering cost of a multi-vendor voice agent stack

Written by

Devon Malloy

Published on

22 September 2026

Your stack works. That’s the part nobody warns you about.

You wired up a speech-to-text vendor, an LLM, a text-to-speech vendor, and an orchestration layer — LiveKit or Pipecat, most likely. You tuned the latency budget. You got the barge-in behavior feeling right. The demo lands, the pilot ships, and the architecture diagram looks reasonable on a whiteboard.

Then it’s nine months later and you have an engineer whose actual job, if you’re honest about how her week goes, is keeping four vendors in agreement with each other.

That’s the cost this post is about. Not whether a cascaded stack can hit the quality bar — it can, and where it can’t is a separate conversation that the production ceiling post handles properly. This one assumes your stack is good. The question is what it costs you, every week, forever, to keep it that way.

What you’re actually signing up for

Four vendors is the number most teams land on. It’s not four integrations. It’s four of everything, permanently.

Surface What it adds to your team’s load
Four onboarding flows A separate account, a separate set of credentials on a separate rotation schedule, and a separate SDK with its own auth model and its own idea of what a “session” even is. Then multiply that by every environment you run. The onboarding is the cheap part and it’s still the part that eats the first sprint.
Four billing relationships Speech-to-text bills per audio minute. The LLM bills per token. Text-to-speech bills per character. The orchestration layer bills per session, or per concurrent connection, or per participant-minute. Four units that don’t reconcile into one another, which means nobody on your team can answer “what does a call cost?” without opening a spreadsheet — and nobody can answer “what will a call cost at 10x?” at all.
Four observability contexts When a turn goes wrong, the evidence for it is split across four dashboards, with four clock skews, four retention windows, and four ideas of what a request id is. There is no single trace for the turn. Reconstructing one is a manual join you do at 2am, by hand, under pressure.
Four failure surfaces Your availability is the product of your vendors’ availability, not the best of them. Four status pages to watch, four incident channels to sit in, four escalation paths with four different response-time commitments. Any one of them can take your agent down by itself, and the one that does will be the one you hadn’t thought about.
Four release cycles Every model update, deprecation notice, and breaking change lands on your roadmap whether or not you asked for it. Four vendors shipping independently means a regression can show up in production without a single line changing in your repository — which is a genuinely awful class of bug to debug, because your first instinct is to look at your own diff.

That’s four onboarding flows, four credential rotations, four billing models, four dashboards, and four release calendars before you write a meaningful line of product code. None of it is your product. All of it is yours to maintain.

Start On The Free Tier

Run one of your own call recordings through a single unified session before you scope the four-vendor version. An afternoon of that is worth more than a week of architecture diagrams, because the thing you’re really trying to price is the maintenance, and you can only feel that by building both.

Sign up free

The operational burden is the architecture

Here’s the part I want to be precise about, because it’s easy to read the table above as a list of things a sufficiently disciplined team would just handle.

It isn’t. The burden isn’t a consequence of doing the integration badly. It’s a consequence of the integration existing.

Every vendor boundary in a voice pipeline is an interface you own forever, and voice interfaces are unusually unforgiving. A turn has to cross all four boundaries inside a budget the caller perceives as roughly a second. Each boundary adds serialization, a network hop, and a queue. More to the point, each boundary is a place where two vendors have to agree on something neither of them coordinates with the other about: what counts as end of turn, what an interruption means, whether partial transcripts should be forwarded or buffered, how a barge-in cancels in-flight synthesis.

Nobody owns the turn. That’s the actual architectural problem. Your speech-to-text vendor owns transcription. Your LLM owns the response. Your text-to-speech vendor owns the audio. Your orchestration layer owns the plumbing. The turn — the unit your users actually experience, and the unit your quality metric is defined on — is owned by no one but you.

Which is why the observability row above is the one that quietly costs the most. You can’t instrument a thing that doesn’t exist as a first-class object in any of your vendors’ systems. So you build the correlation yourself: a trace id you thread through four SDKs that weren’t designed to carry one, a log shipper that normalizes four schemas, and a dashboard that stitches them into something a human can read during an incident. That’s a real internal product. It has a maintainer. It breaks when any of the four vendors changes a field name.

And every latency regression becomes a negotiation. Something got 200ms slower. It’s in one of four systems, measured four different ways, and the only way to find out which is to go looking. Meanwhile the caller just hears a pause.

What “collapsed” actually means

The alternative is collapsing the pipeline into one vendor and one connection. Worth being concrete about what that does and doesn’t mean, because “unified” is a word that gets used loosely.

The Voice Agent API is one WebSocket. Speech recognition, LLM reasoning, and voice generation happen behind it, on Universal-3.5 Pro Realtime. It’s JSON over a socket, and there’s no SDK required to use it.

Run that back through the five rows. One account, one set of credentials. One bill, in one unit, at one rate. One event stream, so the turn exists as an object you can actually query — session recordings and transcripts are retrievable after the fact without you having built the correlation layer yourself. One availability number instead of a product of four. One release cycle to track.

The surface area difference is the part that surprises people. The Voice Agent API exposes 7 client-to-server and 14 server-to-client event types. That’s the whole protocol, and it’s in the events reference. Compare that to the OpenAI Realtime API’s 30+ event types — and then remember that in a four-vendor stack, the event surface you’re actually holding in your head is the union of four separate ones, each with its own conventions.

That’s the trade. You give up granular swap-ability. You can’t drop in a different text-to-speech voice vendor next quarter because a benchmark moved. In exchange, the thing you’re maintaining is one interface instead of six, and the on-call rotation has one place to look.

Whether that trade is right for you depends on things I can’t know from here. The voice agents guide walks through the tradeoffs in more detail than fits in a blog post, and the cluster hub post has the benchmark numbers if quality is the axis you’re evaluating on.

The math, actually

Two numbers, and they point in the same direction.

Availability compounds downward. Suppose each of your four vendors independently hits 99.9%. That’s a good number — most of them publish it. Your agent’s availability is not 99.9%. It’s 0.999 to the fourth power, which is 99.6%.

In hours: 99.9% is about 8.8 hours of downtime a year. 99.6% is about 35. You didn’t do anything wrong to earn the extra 26 hours. You bought them the moment you drew a box diagram with four boxes in series. And that’s the optimistic version, because it assumes the four failures are independent, which they aren’t — they share cloud regions, they share upstream providers, and they occasionally share a bad afternoon.

Cost gets legible instead of estimated. The Voice Agent API is $4.50/hr flat, billed per second of session duration, with speech recognition, reasoning, and voice generation in that one line item. So:

A 5-minute call is $0.375. 100 concurrent calls running for 8 hours is $3,600. 1,000 calls a day at 4 minutes each is $300 a day.

You can do that arithmetic in your head, in a meeting, in front of a CFO. Try that with a four-vendor stack where one component bills per minute, one per token, one per character, and one per concurrent connection. You’ll get a range, and the range will be wide, and the wide part will be the LLM because token counts move with prompt length and nobody’s prompt length is stable. The pricing page has the full breakdown.

But the number neither model captures is the engineering one. An engineer maintaining vendor integrations full-time costs more than any of this. Four vendors doesn’t necessarily mean four engineers — it usually means one or two people whose calendars never come back. That’s the line item that doesn’t appear on any invoice, which is exactly why it doesn’t get budgeted, which is exactly why it’s a surprise nine months in.

Try It In The Playground

Run audio from your actual deployment through it before you take any of this arithmetic on faith. Your traffic shape, your call lengths, and your concurrency pattern determine whether the numbers above are conservative or generous for you.

Try playground

This isn’t an argument against complexity

I want to be straight about where this argument stops, because a post like this can slide into “unified good, cascaded bad” and that would be wrong.

Some teams should absolutely run four vendors. If you’re doing something genuinely novel at one layer — a custom voice you trained, a reasoning setup that needs a specific model with a specific tool-calling behavior, a routing layer with logic no vendor would build for you — then the boundary is where your differentiation lives. Paying the integration tax to own that layer is a good trade. That’s not overhead, that’s the product.

Same if you have real scale and real negotiating power. At high enough volume, negotiating four contracts separately beats one bundled rate, and you likely have the platform team to absorb the maintenance. The burden doesn’t disappear, it just gets amortized across enough revenue to stop mattering.

And if you’re genuinely uncertain about a layer — you don’t know yet whether your callers will tolerate a particular voice, or which LLM handles your domain — keeping that boundary open has option value. Optionality costs something. Sometimes it’s worth what it costs.

What I’d push back on is paying the tax by default. Most teams end up with four vendors because that’s how the tutorials were written, not because they made a decision. The stack accretes. Nobody sits down and says “we’re committing to permanent multi-vendor maintenance in exchange for swap-ability at three layers we will never swap.”

For what it’s worth, we don’t treat the cascaded path as a competitor. Universal-3.5 Pro Realtime ships as a first-class integration on both LiveKit and Pipecat, and the Voice Agent API lets you connect your own LLM if the reasoning layer is where your differentiation sits. The choice isn’t binary. You can collapse three boundaries and keep the one that matters.

The real decision variable is headcount

Here’s what I’d actually change about how these decisions get made.

Stack architecture gets evaluated as a technical question — latency, accuracy, flexibility — and then it quietly becomes a staffing question that nobody revisits. The four-vendor stack isn’t a design choice you make once. It’s a standing headcount commitment you make once and then renew silently, every quarter, in the form of someone’s calendar.

So price it as one. When you’re scoping the architecture, write down the name of the person who will own the vendor boundaries in eighteen months, and what else they won’t be doing. If that name is easy to write down and the tradeoff is obviously worth it, run four vendors — you’ve made an actual decision. If you can’t write down the name, you haven’t priced the stack. You’ve priced the build.

The build is the cheap part. It always was.

Weighing A Multi-Vendor Build Against A Collapsed One

The cost modeling in particular goes faster in a conversation than in a spreadsheet, because the variables that matter are your concurrency pattern and your call length distribution, and those are quicker to say out loud than to document.

Talk to AI expert

Frequently asked questions

What is a multi-vendor voice agent stack?

It’s a voice agent assembled from separately contracted components, most commonly four: a speech-to-text vendor, an LLM, a text-to-speech vendor, and an orchestration framework such as LiveKit or Pipecat that routes audio and events between them. It’s also called a cascaded stack, because audio cascades through each stage in sequence. The architecture is flexible by design — you can swap any layer independently — and the cost of that flexibility is that every boundary between layers is an interface your team owns and maintains permanently.

How does a collapsed voice agent stack work?

A collapsed stack replaces the vendor boundaries with a single connection. On the AssemblyAI Voice Agent API, you open one WebSocket and speech recognition, LLM reasoning, and voice generation all happen behind it, on Universal-3.5 Pro Realtime. You send audio and configuration up; you get transcripts, agent responses, and synthesized audio back down. The protocol is 7 client-to-server and 14 server-to-client event types, in JSON, with no SDK required. Because the full turn happens inside one system, the turn exists as something you can query after the fact rather than something you have to reconstruct by correlating four sets of logs.

When should a team use a multi-vendor voice stack instead of a unified one?

Three cases make it a clear win. First, when a specific layer is your differentiation — a custom-trained voice, an unusual reasoning setup, routing logic no vendor would build for you — because then the boundary is where your product lives rather than where your overhead lives. Second, when you have enough volume that negotiating four contracts separately beats one bundled rate, and enough platform engineering to absorb the maintenance without it showing up on the product roadmap. Third, when you’re genuinely uncertain about a layer and the option to swap it has real value to you. Outside those cases, most teams end up multi-vendor by accretion rather than by decision, which is the situation worth re-examining.

How much does the AssemblyAI Voice Agent API cost?

$4.50 per hour, flat, billed per second of session duration. That single rate covers speech recognition, LLM reasoning, and voice generation — there are no per-layer add-ons. In practical terms: a 5-minute call costs $0.375, 100 concurrent calls running for 8 hours costs $3,600, and 1,000 calls a day at 4 minutes each costs $300 a day. The value of a flat per-second rate isn’t only that it’s cheap, it’s that it’s predictable — you can forecast a month of traffic without modeling token counts or character counts.

What’s the difference between the Voice Agent API and the OpenAI Realtime API?

The clearest practical difference is protocol surface area. The Voice Agent API exposes 7 client-to-server and 14 server-to-client event types; the OpenAI Realtime API exposes 30 or more. That difference compounds across a team, because every event type is something someone has to understand, handle, and eventually debug at an inconvenient hour. The Voice Agent API also runs on Universal-3.5 Pro Realtime for the speech recognition layer, bills at a flat $4.50/hr per second of session, and supports connecting your own LLM if you want to keep the reasoning layer under your control.

How do I debug a failing turn in a unified voice pipeline?

Start from the event stream rather than from the audio, because the event stream is already ordered and already correlated to the session. Pull the session’s history to get the transcripts and the recording together, then walk the server-to-client events in sequence to find where the turn diverged from what you expected — a turn that ended early, a transcript that resolved differently than the caller intended, a tool call that returned late. The advantage over a four-vendor stack is that this is a single query against a single system with a single clock, rather than a manual join across four dashboards with four retention windows. The troubleshooting guide covers the common failure signatures and what each one usually means.