Insights & Use Cases
September 30, 2026

The voice agent launch runbook: ramp, rollback, and on-call

An operating model for after launch: staged traffic ramps, rollback thresholds, severity definitions, on-call ownership, and 10x traffic spikes.

Kelsey Foster
, 
Growth
Reviewed by
No items found.
Table of contents

Most voice agent advice ends at the moment you flip the switch. Load test it, instrument it, ship it — and then the article stops, right at the point where your problems start.

The launch isn't the hard part. The hard part is week one, when 3% of calls are going somewhere strange and you have to decide, at 2am, whether that's a blip or a reason to pull a model version. Nobody writes that part down until they've been burned by it.

So here's the operating model for after launch: how to ramp traffic, what to write in a rollback runbook before you need one, how to define severity for a system whose failures are conversational rather than binary, and what to expect from your own on-call rotation versus your vendor's.

If you haven't launched yet, start with load testing your agent before launch and what to instrument when you go live. This picks up where those leave off.

Why voice agents break differently

A normal web service fails loudly. The request 500s, the error rate spikes, the dashboard turns red, someone gets paged.

A voice agent mostly fails quietly. The call connects. Audio flows. The transcript comes back. Every metric you'd normally watch looks fine. And the caller is repeating their account number for the third time because the agent keeps hearing it wrong.

That's the whole problem in one sentence. Your uptime is 100% and your product is broken.

Three properties make voice different from the systems your on-call process was probably designed around:

  • Failures are partial and conversational. An agent doesn't fail — a turn fails. The call continues, degraded, and the damage accumulates across the conversation rather than announcing itself in one event.
  • Latency is experienced, not measured. A 400ms P99 sounds fine on a chart. In a conversation it's a pause the caller notices, and two of them in a row reads as "this thing is broken."
  • Cost scales with conversation length, not request count. A degraded agent makes calls longer. So a quality regression and a cost incident are frequently the same event, discovered in different dashboards by different teams a week apart.

Each of these has a consequence for how you ramp, when you roll back, and what you page on. Let's take them in order.

Build on speech accuracy that holds up in production

The Voice Agent API runs on Universal-3.5 Pro Realtime — one WebSocket instead of three providers, at a flat $4.50/hr. Start with a free account and clear docs.

Sign up free

The traffic ramp: canary percentages and hold points

Don't go from 0% to 100%. That's obvious. What's less obvious is that the standard software canary — 1%, 10%, 50%, 100%, each for fifteen minutes — is wrong for voice, because fifteen minutes doesn't contain enough conversations to tell you anything.

Voice traffic is heterogeneous in ways web traffic isn't. Morning callers are different from evening callers. Monday is different from Saturday. The angry-customer cohort calls at different times than the quick-question cohort. A canary that only sees one slice of your day tells you your agent works for that slice.

Hold each stage long enough to cover a representative mix, and gate on conversation counts rather than clock time:

StageTrafficHold untilWhat you're actually checking
Shadow0% live, 100% mirroredA few hundred conversationsDoes the new version disagree with the old one, and where
Canary1–2%A full business dayCatastrophic regressions, integration breaks
Early10%A full week, including a weekendCohort-specific failures, containment rate, handle time
Majority50%A few daysCost per conversation, escalation volume, capacity
Full100%Keep the old version warm for 14 daysSlow-burn regressions that only show at volume

Two things to get right here.

Route by conversation, not by request. If a single call can hit both the old and new version mid-conversation, your canary data is garbage and so is the caller's experience. Pin the version at session start and hold it for the life of the call.

Keep a control group. Don't ramp the last 10% for a while. Without a live baseline running the same day against the same caller mix, you can't distinguish "the new version got worse" from "Tuesday was worse." This is the single most common thing teams skip, and it's the thing that makes every subsequent argument unresolvable.

Shadow mode deserves a note. Mirroring live audio to a candidate version and diffing the transcripts is the cheapest signal you will ever get, because it costs you nothing in caller experience. The disagreements between versions are where the interesting cases live — not the average WER, but the specific turns where two models saw the same audio and reached different conclusions.

One planning detail that bites here: a shadow stage doubles the number of streaming sessions you open per minute, and new-session rate is the thing that's actually governed (see the spike section below). Factor that in before you turn mirroring on at 100%.

Writing the rollback runbook before you need it

At 2am, nobody makes good judgment calls. They execute decisions that were made earlier by people who weren't tired. That's the entire purpose of a rollback runbook: to move the decision from the incident to a Tuesday afternoon when everyone was calm.

A runbook that says "roll back if quality degrades" is not a runbook. It's a wish. The thresholds have to be numbers, and the numbers have to be tied to metrics you're actually collecting.

Pick thresholds you can defend

Set these per agent, based on your own baseline — the specific numbers below are illustrative, not universal:

  • Containment rate drop. The share of conversations resolved without a human. If it falls more than X% below the control group over a rolling window of N conversations, roll back. This is the best single indicator because it captures quality through outcome rather than through a proxy.
  • Escalation-to-human spike. Same signal, different direction, and usually faster to move.
  • Repeat-utterance rate. How often a caller says substantially the same thing twice in a row. This is your best early proxy for "the agent isn't understanding," and it moves before containment does.
  • Turn latency at P90, not P50. Your median can look great while a tenth of turns feel broken. Alert on the tail.
  • Average conversation duration. Rising duration with flat containment means the agent is working harder for the same result — a regression that no error rate will show you.
  • Cost per resolved conversation. The business version of the previous line, and the one your finance team will notice first.

Decide who pulls the trigger

Write down, by name or by role, who can roll back without asking permission. Then make the bar for rolling back low and the bar for rolling forward high. Rolling back a voice agent version costs you a few days of progress. Not rolling back costs you calls you can't get back and a support queue you'll be digging out of for a week.

The runbook should fit on one page and answer five questions: what triggers a rollback, who can call it, what the exact command or config change is, who gets notified, and what you capture before you revert. That last one matters — if you roll back without saving the failing audio, transcripts, and session IDs, you've fixed the symptom and thrown away the diagnosis.

Keep the old version pinnable

A rollback plan only works if there's something to roll back to. Pin an explicit model version in production config rather than always tracking latest, so "revert" is a config change and not a negotiation. On AssemblyAI, previous flagship models stay available as pinnable identifiers for exactly this reason — you can name a specific model rather than taking the current default and get deterministic behavior while you investigate.

Auto-upgrade is a great default for a dev environment. It is a strange default for the system that answers your phone.

Test a rollback candidate on your own audio

Run real call recordings through streaming transcription and compare versions side by side before you commit to a ramp.

Try playground

Severity definitions for a system that fails conversationally

Standard severity ladders are built around availability. SEV1 is "it's down." For a voice agent, "it's down" is actually one of the easier failures — it's loud, it's obvious, and your existing alerting already catches it.

The failures that hurt are the ones that don't trip an availability alarm. Your severity definitions need to name them explicitly, or on-call will keep classifying real incidents as "monitoring noise."

SeverityDefinitionExampleResponse
SEV1Calls are not being served, or are being served wrong in a way that causes real-world harmConnection failures; the agent confirms actions it did not take; PII spoken back to the wrong callerPage immediately, fail over to human queue, roll back without discussion
SEV2The agent is up but materially degraded for a large share of callersContainment below rollback threshold; P90 turn latency past the conversational limit; repeat-utterance rate doubledPage during business hours, roll back if it persists past one window
SEV3A specific cohort or intent is broken while overall metrics look normalOne accent or language degraded; a single intent failing; spelled-out identifiers failingTicket, investigate next business day, consider partial rollback
SEV4Cost or efficiency regression with no quality impactConversation duration up with containment flatTrack it; fold into the next iteration

SEV3 is the row most teams don't have, and it's the row that catches the failures that actually erode trust. A cohort-specific regression is invisible in aggregate — if 6% of your callers are having a terrible time and the other 94% are fine, every dashboard you own will look green. Segment your quality metrics by language, by accent where you can infer it, by intent, and by channel. If you can't slice it, you can't see it.

On-call: what you own and what your vendor owns

Voice agents have more moving parts than most teams' incident processes assume. Telephony, speech-to-text, the LLM, text-to-speech, your own application logic, and whatever CRM or database the agent reaches into. When a call goes wrong, the first job isn't fixing it — it's figuring out which of those six things caused it.

This is the strongest practical argument for consolidating the pipeline. With separate STT, LLM, and TTS providers you have three vendors, three sets of logs, three status pages, and three support queues, and the boundary between them is exactly where the hard bugs live. Running the Voice Agent API gives you one WebSocket, one bill, and one set of logs, which collapses the triage step from "which vendor" to "what happened."

Whatever your architecture, write down the boundary explicitly:

  • You own your application logic, your prompts and agent configuration, your ramp and rollback tooling, your quality metrics and their thresholds, your fallback path to humans, and the decision to pull a version.
  • Your vendor owns model availability, published latency characteristics, capacity and rate limits, incident communication, and — if you've negotiated it — a response time when you're the one paging them.

What to ask a vendor before you need them at 2am

Ask these in evaluation, not during an incident:

  1. Is support actually 24/7, and what's the response target for a production-down issue? "We have a support email" is not an answer.
  2. Can I reach a human who understands the model, or only a ticket queue? The distinction matters enormously at 3am.
  3. How are model versions deprecated, and how much notice do I get? A silent upgrade under your production traffic is an incident someone else scheduled for you.
  4. Can I pin a version? If not, you have no rollback target.
  5. Where's the status page, and what actually gets posted to it? Then go read its history. A status page with no incidents on it usually means the bar for posting is high, not that nothing has ever broken.
  6. What exactly is rate limited, what are my current numbers, and what happens when I exceed them? "We scale automatically" is the beginning of the answer, not the end. Get the unit — sessions, requests, concurrent jobs — and get the failure mode.

That last one is worth dwelling on, because it's where the cost question and the reliability question meet.

The 10x spike is a capacity problem and a cost problem at the same time

Volume spikes aren't hypothetical. A product launch, an outage at a company you integrate with, a news cycle, a billing run — all of these can multiply call volume overnight, and the spike arrives as two simultaneous incidents.

The capacity side. Find out now, not later, what your provider limits and what it does when you cross the line. Treat any vendor's "unlimited" as a prompt to ask what the unit is, because the constraint is usually real and just measured in something other than what you assumed.

For AssemblyAI streaming specifically, the shape is worth knowing because it's a good fit for spikes and a bad fit for cold starts:

  • There is no hard cap on the number of sessions you can have open at once. That's the part people mean when they say "unlimited concurrency," and for steady-state load it's accurate.
  • What is limited is how many new sessions you can open per minute — 5 on free accounts, 100+ on paid.
  • That limit auto-scales with no ceiling: any minute you use 70% or more of your current allowance, next minute's allowance goes up 10%. Sustained, that compounds fast — roughly 100 → 110 → 121 → 133 → 146 across five minutes, and it keeps going.
  • Exceed the current allowance and new connections are rejected with WebSocket close code 1008, "Too many concurrent sessions." Sessions already open are unaffected.
  • It also scales back down when your usage drops below 50% of the limit.

Read those last two together and the operational risk is specific: a 10x spike that arrives after a quiet period hits a limit that has already shrunk toward your account's baseline, and climbing back is a 10%-per-minute ramp. That's the scenario to plan for. If you know a spike is coming — a launch, a campaign, a migration — request a higher limit in advance, at no additional cost. If you don't know, build a retry with backoff on 1008 and degrade gracefully rather than dropping the caller.

Pre-recorded transcription is governed differently and more gently: the limit is the number of jobs processing in parallel (5 free, 200+ paid), and anything over it is queued FIFO rather than rejected. So a backlog costs you latency, not calls. There's a separate HTTP ceiling of 20,000 requests per five minutes across all endpoints, which you hit by polling too aggressively — use webhooks, or widen and jitter your polling interval.

Ramp behavior at real scale is worth planning for specifically. Contact center and CX platforms routinely bring millions of hours of monthly volume online in a matter of days rather than months, which is exactly the pattern the auto-scaling is designed for — but "designed for" still means telling someone first when you can.

The cost side. Per-component billing makes spikes hard to forecast, because your bill is a function of conversation length, model calls per turn, and token counts, and all three move together when things go wrong. Flat-rate pricing — the Voice Agent API is $4.50/hr all in — turns cost forecasting into multiplication. Whichever model you're on, the operational requirement is the same: a cost alert tied to conversations per hour, set at a multiple of your normal peak, firing to the same rotation as your quality alerts. A cost incident found in next month's invoice is a cost incident you paid for in full.

The one-page runbook

Everything above, compressed to the thing you actually pin in the channel:

FieldFill in before launch
Current version / previous versionExplicit pinned identifiers, not "latest"
Rollback commandThe exact config change or deploy command, copy-pasteable
Rollback authorityNamed roles who can act without escalating
Rollback thresholdsContainment, escalation, P90 latency, repeat-utterance — with numbers
Capture-before-revertSession IDs, audio, transcripts, config snapshot
Human fallbackHow to route 100% to people, and who staffs it
Vendor escalationContact path, response target, status page URL
Rate limitsCurrent new-sessions/min, who to ask for a raise, behaviour on 1008
Cost circuit breakerConversations/hour threshold and who it pages

Reliability is a feature your customers can see

There's a version of this post that's purely operational, and there's a more honest one, which is this: your reliability is visible to the people who buy from you.

Status pages are public. Incident histories are public. When a voice infrastructure provider has a rough month, their customers notice, their customers' customers notice, and it becomes part of how the market talks about them. The same is true of the agent you're about to launch — your callers don't distinguish between "our vendor had an incident" and "this company's phone system doesn't work."

Which reframes the whole exercise. A ramp plan, a rollback runbook, and a severity ladder aren't bureaucracy you add because someone made you. They're the mechanism by which the reliability you're selling stays true on the days when something breaks. And something will break — the difference between teams is only whether the decision about what to do next was made in advance or at 2am.

Write it down this week. You'll need it sooner than you think.

Ship a voice agent you can actually operate

One WebSocket, flat $4.50/hr, auto-scaling session limits with no ceiling, and 24/7 support from engineers who know the models. Get a free API key and start building.

Sign up free

Frequently asked questions

What's the rollback plan if a new voice agent model version underperforms in production?

Pin an explicit previous model version in production config so reverting is a single config change, and define numeric rollback thresholds before you ramp — typically a containment-rate drop versus a live control group, an escalation-rate spike, and P90 turn latency past your conversational limit. Name the roles who can execute a rollback without escalating, and capture session IDs, audio, and transcripts before reverting so you keep the diagnosis. Rolling back should be cheap and rolling forward should require evidence, not the other way around.

How do I monitor latency and call quality for a live voice agent in production?

Track turn latency at P90 rather than P50, because the tail is what callers experience as a broken pause, and pair it with outcome metrics — containment rate, escalation rate, average conversation duration, and repeat-utterance rate — segmented by language, intent, and channel. Aggregate dashboards hide cohort-specific failures, so a regression affecting one accent or one intent will look green until you slice the data. Keep a live control group on the previous version so you can tell a real regression apart from a bad Tuesday.

What happens to my costs if my voice agent call volume spikes 10x overnight?

With per-component STT + LLM + TTS billing, a 10x call spike is more than a 10x cost increase, because degraded agents produce longer conversations and more model calls per turn. Flat-rate pricing makes the math linear — AssemblyAI's Voice Agent API is $4.50/hr all in. On the capacity side, streaming puts no cap on how many sessions you can hold open, but it does limit how many new sessions you can open per minute (100+ on paid accounts), auto-scaling by 10% each minute you run at 70% or more of that allowance, with no ceiling. Exceeding it closes new connections with code 1008 while existing calls continue. Because the limit also scales back down after quiet periods, a spike arriving cold is the case to plan for: request a higher limit ahead of a known event, and retry with backoff otherwise. Either way, set a cost circuit breaker on conversations per hour and route it to the same rotation as your quality alerts.

What support is available if my voice agent has issues in production at 2am?

Ask before you launch, not during an incident: whether support is genuinely 24/7, what the response target is for a production-down issue, and whether you reach an engineer who understands the models or only a ticket queue. AssemblyAI provides 24/7 support and forward-deployed engineers who work as embedded members of customer teams. Regardless of vendor, your own runbook should define a human fallback path so you can route calls away from the agent while you wait.

How long should I run a canary before ramping a voice agent to full traffic?

Gate on conversation counts rather than elapsed time, and hold long enough to see a representative caller mix — a canary that only sees Tuesday morning tells you your agent works on Tuesday mornings. A practical ladder is shadow mode for a few hundred conversations, 1–2% for a full business day, 10% for a week including a weekend, then 50%, keeping the previous version warm for two weeks after full rollout. Pin the version for the life of each call so a single conversation never spans two versions.

What's the difference between a voice agent being down and being degraded?

Down is loud and your existing availability alerting already catches it; degraded is silent and is where the real damage happens. A degraded agent connects the call, returns a transcript, and keeps every infrastructure metric green while the caller repeats themselves four times. Severity definitions for voice need an explicit tier for "up but materially degraded" and another for "one cohort is broken while aggregates look normal," or on-call will keep classifying genuine incidents as monitoring noise.

Title goes here

Lorem ipsum dolor sit amet, consectetur adipiscing elit, sed do eiusmod tempor incididunt ut labore et dolore magna aliqua. Ut enim ad minim veniam, quis nostrud exercitation ullamco laboris nisi ut aliquip ex ea commodo consequat. Duis aute irure dolor in reprehenderit in voluptate velit esse cillum dolore eu fugiat nulla pariatur.

Button Text
AI voice agents
Voice Agent API