Voice AI for transcription services: the infrastructure underneath
Learn how to use Speech-to-Text and Voice AI technology to transform your transcription services and upgrade your offerings.



If you run a transcription service, Voice AI isn't a product you buy. It's the layer your product sits on — and the interesting question isn't what it can do, it's what it costs you per hour and how often it gets a name wrong.
What is Voice AI, in the part that matters here?
Voice AI is the set of models that turn spoken audio into text and then into structured meaning. For a transcription business, that breaks into three things you actually buy: recognition accuracy, the languages you're covered for, and the price per hour of audio.
Everything downstream — summaries, search, speaker labels, sentiment, redaction — depends on the transcript being right first. A summary built on a bad transcript is a confident, well-formatted mistake. We've written more about that gap in transcription accuracy versus transcription quality.
Which model should a transcription service run?
Three, depending on the job. Here's the current line, with the real prices.
Universal-3.5 Pro averages 4.35% normalized word error rate across our benchmark datasets, and the realtime version averages 5.53%. Both figures, plus the diarization and code-switching comparisons, are on our benchmarks page. The Pro line also produces the transcript and the speaker changes jointly, which is why it catches short turns and overlapped speech that a bolted-on diarization pass tends to smear.
You don't have to pick one forever. Set speech_models: ["universal-3-5-pro", "universal-2"] and you get the flagship with a documented fallback. Or omit the field entirely and auto-upgrade to whatever the latest Universal Pro model is.
The transcript is the input, not the product
This is where most transcription services find their margin. Once audio is text, the Speech Understanding API returns entity detection across 50+ types, sentiment, topic detection, key phrases, content safety flags, PII redaction in both text and audio, translation into 100+ target languages, and audio event tagging — all from the same request, without you standing up a second pipeline. That's the same stack behind most conversation intelligence products.
For anything that needs a language model on top — custom summaries, structured extraction, chaptering, question answering over a transcript — the LLM Gateway gives you one API to OpenAI, Anthropic, Google and others. One key, one bill, and no vendor migration when you want to try a different model next quarter.
That's the whole shape of it: speech-to-text as a primitive, understanding as a primitive, generation as a primitive. You compose them. We don't have opinions about what you build.
What people have built on it
Two customers put the value more precisely than we can.
"AssemblyAI has a real high-touch personal service. It's a great partnership and we're very collaborative and get to test new AI models early and work together. And AssemblyAI is really pushing boundaries, helping us create a well-rounded conversation intelligence platform." — Tom Lavery, CEO & Founder, Jiminny
"The cost saving is literally the difference between being profitable or not for us, but beyond the economics, AssemblyAI gave us something invaluable: peace of mind. We can focus on building our product instead of worrying about infrastructure limits." — Mark Barbir, CEO, Earmark
Earmark's point is the one transcription businesses underrate. At scale, the per-hour rate isn't a line item — it's the difference between a viable unit economic and a treadmill. And there's demand-side evidence too: in our 2025 INSIGHTS REPORT: The State of Conversation Intelligence, 69% of companies cited improved customer service after implementing conversation intelligence.
Getting started
Sign up, get a key, and send a file. Pricing is pay-as-you-go and billed per second with no minimums, so you can test at real volume before committing to anything. The pricing page has the full table including add-ons, and the docs cover the request shapes for every model above.
One last thing worth saying plainly. The reason to think of this as infrastructure rather than a transcription product is that the models keep changing underneath you — three flagship releases in the last year alone. If you've built against a platform, each of those is a migration. If you've built against an API, most of them are a one-line change, or nothing at all.
Frequently asked questions
What features does AssemblyAI offer for meeting transcription?
Speaker diarization is the core one — Universal-3.5 Pro produces the transcript and speaker changes jointly, which handles rapid back-and-forth and overlapped speech better than a separate clustering pass. On top of that you get entity detection, topic detection, key phrases, and summaries through the Speech Understanding API and LLM Gateway. Contextual prompting lets you prime the model with the meeting agenda or participant names before the file is processed.
How does AssemblyAI handle large-scale enterprise conversation processing?
Pay-as-you-go billing per second with no minimums, so throughput scales without a contract renegotiation. Deployment options cover our hosted Voice AI Cloud, self-hosting inside your own AWS VPC, and EU data residency at the same price with data staying in the EU. For organizations subject to HIPAA, AssemblyAI is considered a business associate and offers a standard Business Associate Addendum, which can be signed in minutes without a sales call.
How does AssemblyAI handle both live and recorded calls?
They're separate endpoints sharing a model family. Recorded audio goes to the async API on Universal-3.5 Pro at $0.21/hr, and live audio goes to the streaming WebSocket on Universal-3.5 Pro Realtime at $0.45/hr, billed on session duration rather than audio length. There's also a Sync API for short clips that returns a transcript in one request with no polling.
What's the most accurate API for transcribing YouTube and video content?
Universal-3.5 Pro averages 4.35% normalized word error rate across our benchmark datasets, which is the strongest result we've published for pre-recorded audio. For video specifically, the things that usually break transcription are music beds, multiple speakers, and code-switching mid-sentence — all of which the Pro line handles natively across its 18 languages. Check the public benchmarks page for the per-dataset breakdown rather than taking the average on faith.
How do you generate subtitles for video at scale?
Submit files to the async API and request subtitle output alongside the transcript; billing is per second of audio, so a large back catalogue processes in parallel. Universal-2 at $0.15/hr is the better choice when your library spans languages outside the Pro line's 18, since it covers 99+. Translation into 100+ target languages is available on pre-recorded audio for multilingual subtitle tracks.
How much does high-volume transcription actually cost?
Universal-3.5 Pro is $0.21 per hour of audio and Universal-2 is $0.15, both billed per second with no minimum spend. Add-ons are priced separately and individually — standard diarization is $0.02/hr, keyterms prompting is $0.05/hr — so you only pay for the ones you turn on. There's no seat licensing and no platform fee on top.
Lorem ipsum dolor sit amet, consectetur adipiscing elit, sed do eiusmod tempor incididunt ut labore et dolore magna aliqua. Ut enim ad minim veniam, quis nostrud exercitation ullamco laboris nisi ut aliquip ex ea commodo consequat. Duis aute irure dolor in reprehenderit in voluptate velit esse cillum dolore eu fugiat nulla pariatur.



