When to stop self-hosting Whisper (and what you actually gain)
Self hosting Whisper gets expensive fast. See the real GPU, DevOps, and scaling costs, plus when a managed speech-to-text API makes more sense for many teams.



Self-hosting Whisper is free to license and expensive to run. Between GPU instances, initial setup, ongoing patching, and capacity planning, most teams below high, steady volume spend more running Whisper than they would on a managed API—and they still have to build diarization, streaming, and redaction themselves. Here is where the line actually falls.
The fork is simple to state and expensive to get wrong. Managed API, or your own GPUs running an open-source checkpoint.
What makes it hard is that the two options don't fail in the same place. A managed API fails on your invoice, visibly, every month. Self-hosting fails in a hundred small places you never budgeted for: the driver upgrade that broke inference on a Friday, the diarization model you bolted on and now maintain, the engineer who can't be reassigned because she's the only one who understands the queue. We'll compare both across the factors that actually decide it, with numbers where we have them, and we'll be straight about the cases where self-hosting is the right call. There are several.
AssemblyAI vs Whisper: at a glance
| Aspect | AssemblyAI | Whisper |
|---|---|---|
| Deployment model | Cloud API (managed), EU data residency, plus a self-hosted streaming deployment | Self-hosted (open-source) |
| Pricing | $0.21/hr of audio on the flagship async model, billed per second | Free software (you pay for servers) |
| Key strengths | Built-in features, no maintenance, automatic model upgrades | Complete control, true offline capability, no per-minute bill |
| Best use cases | Production apps, real-time needs, regulated workloads | Research, air-gapped processing, local experimentation |
| Setup complexity | API key plus a few lines of code | GPU setup, model downloads, CUDA, capacity planning |
AssemblyAI and Whisper both convert speech to text, but they work completely differently. AssemblyAI is a cloud-based API service—you send your audio files to their servers and get transcripts back. Whisper is open-source software that you download and run on your own computers.
Think of it like email hosting. You can use Gmail (managed service like AssemblyAI) or run your own email server (self-hosted like Whisper). Each approach has clear trade-offs.
The fundamental question is whether you want convenience or control. AssemblyAI handles everything—updates, scaling, infrastructure—but you depend on their service. Whisper gives you complete control but requires technical expertise to run properly.
One correction to a common assumption before we go further: this is not a strict cloud-versus-on-prem split anymore. AssemblyAI ships a self-hosted deployment that runs inside your own infrastructure for streaming workloads, and it runs in the EU at the same price for teams whose real requirement is data residency rather than air-gapping. More on the exact scope of that below, because the details matter.
Which is more accurate for speech recognition?
Both platforms produce high-quality transcripts. On real-world audio, AssemblyAI's Universal models measure better, and here are the numbers rather than the adjectives.
Universal-3 Pro's hallucination rate is roughly 30% lower than Whisper's—the single most consequential difference, because a hallucination is not a misheard word, it's an invented one that reads as plausible. English mean WER is 5.6% (median 4.9%). On LibriSpeech Clean it posts 1.52% and on LibriSpeech Other 2.69%. CommonVoice comes in at 4.87%.
The current flagship, Universal-3.5 Pro, extends that with native code-switching across 18 languages and the most accurate speaker diarization AssemblyAI has shipped, optimized for cpWER rather than DER. Full figures live on the benchmarks page, and if you want the model-by-model head-to-head rather than the deployment decision, that's covered in depth in AssemblyAI vs Whisper Large-v3.
But here's what's really interesting—Whisper sometimes creates "hallucinations." These are words or phrases that weren't actually spoken but appear in the transcript. They tend to show up in silence, in music, and at the ends of chunks, which is exactly where an automated pipeline is least likely to catch them.
Performance across different audio conditions
Real-world audio isn't perfect. Your users call from noisy cafés, speak with accents, or use cheap microphones. How each platform handles these challenges affects your user experience.
On clean audio, both platforms perform excellently and the differences are minimal. For high-quality single-speaker recordings, the choice comes down to features and implementation, not accuracy.
Challenging audio is where the gap opens, and it opens in four specific places.
- Background noise. Accuracy holds where Whisper degrades, and the streaming line can isolate the primary speaker in near-field or far-field conditions.
- Accents and mixed-language speech—the big one. Universal-3.5 Pro transcribes every word in the language it was actually spoken in, mid-sentence switches included, averaging 7.69 normalized WER on a five-language code-switching benchmark against Deepgram Nova-3 Multilingual at 12.22.
- Want the model to know your product names, your drug names, or your customers' surnames? Hand it up to 1,000 keyterms. Getting the same result out of Whisper means fine-tuning.
- Telephony. AssemblyAI is tuned for it; Whisper wasn't trained with 8kHz call audio as a priority, and you can hear the difference on a bad line.
On multilingual coverage specifically, the honest answer is that AssemblyAI does it two ways. Universal-3.5 Pro handles 18 languages with native code-switching built into the model rather than stitched on afterward, and Universal-2 covers the long tail of 99 languages. Whisper covers 99 languages too, but it has no native code-switching path—you get one language per pass and you build the routing yourself.
Speaker diarization and real-time capabilities
Speaker diarization identifies who's talking when. Instead of a wall of text, you get a conversation with clear speaker labels. This feature transforms meeting transcripts, customer service calls, and interviews from unreadable blocks into usable documents.
With AssemblyAI you send audio and receive a transcript with speaker labels attached, for +$0.02/hr on standard async diarization. With Whisper you transcribe first, run a separate diarization model, then align the two sets of timings yourself and debug the disagreements. That's not a configuration flag, it's a subsystem.
Real-time is the bigger gap. AssemblyAI's streaming speech-to-text runs Universal-3.6 Pro Realtime over wss://streaming.assemblyai.com/v3/ws at $0.45/hr base, with turn detection that combines semantic context with voice activity rather than just silence (the final transcript lands a median 307ms after the speaker stops on Pipecat's open STT benchmark), and diarization-with-revision for up to 10 speakers (+$0.12/hr) that re-clusters and corrects itself within about half a second of the stream ending.
Whisper processes audio in chunks. True real-time is nearly impossible without significant engineering—buffering systems, audio splitting, timing synchronization, and a strategy for what to do when a chunk boundary lands mid-word.
Worth knowing if language coverage is why you're on Whisper: AssemblyAI also offers a managed Whisper-Streaming option covering 99 languages. It isn't listed on the public pricing page, so talk to us for a rate—but the point stands. You don't have to give up Whisper to stop hosting it.
Drop in the file your Whisper deployment gets wrong most often. Check speaker labels, proper nouns, and noise robustness before you change anything.
What features does each platform provide?
| Feature | AssemblyAI | Whisper |
|---|---|---|
| Speaker diarization | Built-in, +$0.02/hr async (standard) | Requires separate models and manual timing alignment |
| Real-time streaming | Universal-3.6 Pro Realtime, $0.45/hr base, final transcript a median 307ms after the speaker stops (Pipecat benchmark) | Batch processing only |
| Sentiment and entity detection | Automatic, via Speech Understanding | Not available |
| Summarization and chapters | Via Speech Understanding and LLM Gateway | Not available |
| PII redaction | Text +$0.08/hr, audio +$0.05/hr | Not available |
| Custom vocabulary | Keyterms prompting, up to 1,000 terms (+$0.05/hr) | Limited support; fine-tuning for real gains |
| Medical accuracy | Medical Mode, one parameter, +$0.15/hr | Requires domain fine-tuning |
The feature gap between these platforms is enormous. AssemblyAI includes many built-in capabilities that would take months to build on top of Whisper.
These aren't just nice-to-have features—they represent significant development work. Building diarization on top of Whisper means integrating additional AI models, handling timing alignment, and debugging when things break. Building Speech Understanding capabilities like sentiment, entity detection, and redaction means an entire second layer of models on top of the transcript you just produced.
How much does it actually cost to self-host Whisper?
| Monthly audio volume | AssemblyAI cost (flagship async, $0.21/hr) | Whisper infrastructure cost |
|---|---|---|
| 1,000 minutes | $3.50 | One cloud GPU instance, billed for the whole month |
| 10,000 minutes | $35 | A dedicated GPU server, still billed for the whole month |
| 100,000 minutes | $350 | Multiple GPUs, plus the engineering time to keep them fed |
Whisper is "free" software, but running it costs money. You need servers with GPUs to process audio quickly, and those GPUs bill by the hour whether or not audio is flowing through them. That's the structural difference: AssemblyAI bills per second of audio, so an idle Saturday costs nothing. A reserved GPU costs the same on Saturday as it does on Monday.
Notice what happens as you scale. The gap narrows, and at high, steady volume it eventually inverts—a fully utilized GPU fleet running a free model is cheaper per hour than any per-second API. That's real, and it's the honest case for self-hosting. The question is whether your utilization is actually high and steady, or whether you're paying for peak capacity to serve a bursty workload. Most teams are doing the second thing.
You are not the only one weighing this. In AssemblyAI's 2026 Voice Agent Report—455 responses fielded across Q4 2025 and Q1 2026—62% of builders named cost-effectiveness as a primary decision driver when choosing speech infrastructure, and 75% said technical reliability was a barrier they were actively fighting. Those two findings pull in opposite directions, which is exactly what makes this decision hard: the cheapest line item and the most reliable system are rarely the same choice.
Hidden costs of self-hosting Whisper:
- Setup time: initial configuration, CUDA drivers, model selection, and benchmarking before you serve a single request
- Maintenance: ongoing updates, security patches, and troubleshooting
- Downtime: when your servers break, your transcription stops
- Scaling challenges: planning capacity for traffic spikes you can't predict
- Engineering resources: DevOps expertise isn't cheap, and it isn't reassignable while it's on call
That last line is the one teams underweight. Here's Mark Barbir, CEO of Earmark, on what changed after they stopped managing their own transcription infrastructure:
"The cost saving is literally the difference between being profitable or not for us, but beyond the economics, AssemblyAI gave us something invaluable: peace of mind. We can focus on building our product instead of worrying about infrastructure limits."
Two arguments in one quote, and they're the two arguments this whole page is about. The line item you can see, and the engineering attention you can't.
For the full model-by-model math on open-source total cost, including GPU utilization curves, we've broken it out separately in the true cost of open-source speech-to-text. Current list prices for every model and add-on are on the pricing page.
What does total cost of ownership actually look like?
The table above compares infrastructure spend, which is the comparison Whisper wins most often. It's also the incomplete one. Here's the full line-item view.
| Cost line | Self-hosted Whisper | AssemblyAI |
|---|---|---|
| Compute | GPU hours, billed idle or busy | $0.21/hr of audio, billed per second |
| Initial setup | Engineering time before first request served | API key |
| Ongoing maintenance | Patching, driver upgrades, on-call rotation | None |
| Diarization | Second model, plus alignment code you own forever | +$0.02/hr async, +$0.12/hr streaming |
| Streaming | Build chunking, buffering, and endpointing yourself | $0.45/hr base |
| PII redaction | Build it | Speech Understanding add-on: +$0.08/hr text, +$0.05/hr audio |
| Model upgrades | Manual: download, benchmark, migrate, re-test | Automatic, or pin a version if you'd rather |
Read that table as a build-versus-buy sheet, because that's what it is. Renting GPUs to run a checkpoint is not the same product as buying a finished per-second API. Four of those seven rows aren't costs at all on the managed side—they're features that already exist.
How long does it take to set up Whisper versus a managed API?
Getting started with each platform reveals the complexity difference immediately.
The AssemblyAI implementation, targeting Python SDK 1.0.0:
# pip install "assemblyai>=1.0.0"
import os
from assemblyai.prerecorded.v2 import Transcriber, TranscriptionConfig
config = TranscriptionConfig(
speech_models=["universal-3-5-pro", "universal-2"],
speaker_labels=True,
)
transcriber = Transcriber(api_key=os.environ["ASSEMBLYAI_API_KEY"])
transcript = transcriber.transcribe(
"https://example.com/interview.mp3",
config=config,
)
print(transcript.text)That's it. A few lines of code and you're transcribing, with speaker labels attached and a fallback to Universal-2 ($0.15/hr, 99 languages) for anything outside the flagship's 18.
Now the Whisper version. First you install CUDA drivers, because GPU acceleration doesn't configure itself. Then several gigabytes of model files come down. Then a Python environment and its dependency graph, which is its own genre of afternoon. Then VRAM budgeting, since the large models want 10GB or more and will tell you about it only at inference time. Then audio preprocessing, because the model is fussier about input than the docs suggest.
None of that is hard. All of it is time, and none of it has transcribed anything yet.
Even after setup, you'll write significantly more code to handle errors, manage processing queues, and implement the features AssemblyAI includes automatically.
Then scaling becomes the real project. AssemblyAI scales by sending more requests. Whisper requires capacity planning: you estimate peak usage, provision for it, and own load balancing, queue management, and failover. When you're building an MVP or trying to ship quickly, that difference sets your timeline.
Swap the transcription call, delete the infrastructure code, keep the diarization and streaming you were about to build. Free API key, no card required.
When Whisper is genuinely the right answer
Most people arriving at this question are not running a GPU fleet. They're running Whisper on a laptop, and for that, Whisper is the correct choice and this entire cost argument is irrelevant.
If you're transcribing your own recordings on a Mac, batching a research corpus overnight, or building something that has to work on a plane, install Whisper and stop reading. It's free, it's good, and there's no bill. The TCO argument only bites when transcription becomes a production dependency other people rely on—when uptime matters, when volume is unpredictable, and when someone has to be paged if it breaks.
One more thing has changed since this post first ran. Self-hosting is no longer a Whisper-only question. Qwen3-ASR, NVIDIA Parakeet and Canary, Voxtral, and Moonshine are all commonly self-hosted now, and most teams doing it run them on Baseten, Modal, or Fireworks rather than on bare GPUs they provision themselves. The economics shift when you're renting per-second inference on someone else's platform instead of reserving instances—we compared those platforms directly in AssemblyAI vs self-hosting on Baseten, Modal, or Fireworks, and we have a separate head-to-head against Qwen3-ASR if that is the model you are weighing.
When to choose AssemblyAI vs Whisper
Your specific situation determines the right choice. Most developers benefit from starting with the managed API, then evaluating alternatives once they understand their actual requirements and volume.
Go managed if any of these describe you. You need to ship in days rather than months. You need real-time—live captions, a voice agent, anything where a chunked batch model simply cannot participate. You want diarization, sentiment, entity detection, and redaction to be parameters instead of projects. Your finance team would like the transcription line to track usage rather than reserved capacity. Or a security reviewer is about to ask what controls you have, and you'd like to answer with SOC 2 Type 2, a pre-signed self-serve Business Associate Addendum for customers processing PHI, PII redaction across audio and transcripts, and EU data residency at api.eu.assemblyai.com at the same price—rather than with a diagram of your VPC.
Choose Whisper when you have:
- True air-gap requirements: audio that can never leave a network with no egress at all, for batch workloads
- Offline needs: no internet connectivity at processing time
- Custom model requirements: fine-tuning on your own data, or modifications to the model itself
- High, steady, predictable volume and the ML engineering capacity to keep a fleet at high utilization
Hybrid approaches work too. Plenty of teams run AssemblyAI for real-time features and difficult audio while running an open model for high-volume batch, abstracted behind a common interface that routes by requirement. That's a legitimate architecture, not a fence-sit.
Final words
Here's the part that doesn't show up in either table. The cost of self-hosting isn't a number that stays where you put it. Every capability you add later—diarization, streaming, redaction, a second language—reopens the build decision, and each reopening is another sprint you didn't plan. A managed API converts those from projects into parameters. That's the actual asymmetry, and it compounds in a direction the spreadsheet doesn't model.
Which reframes the question. It was never really "which is cheaper." It's whether transcription is a thing you're building or a thing you're using—and almost nobody who set out to build it meant to.
Bring your monthly audio hours, your instance types, and your utilization. We'll model the switch honestly, including the volumes where self-hosting still wins.
Frequently asked questions
How much does it cost to self-host Whisper?
Licensing is free; running it is not. Your bill is GPU instance hours, and those accrue whether or not audio is flowing, plus initial setup and ongoing patching, capacity planning, and on-call. Against a managed API billed per second of audio, self-hosting typically only wins at high, steady volume where your GPUs stay busy.
Does AssemblyAI work offline like Whisper?
Partly. AssemblyAI offers a self-hosted deployment that runs Universal-3.5 Pro Streaming inside your own infrastructure, including via the AWS SageMaker Marketplace, so audio never leaves your environment for real-time workloads. Self-hosted async transcription is not currently available, so for fully offline batch processing Whisper remains the option. If the requirement is data residency rather than true offline operation, AssemblyAI also runs in the EU at the same price. Note that the self-hosted deployment does not include Medical Mode, PII redaction, voice focus, or speaker diarization.
Is AssemblyAI more accurate than Whisper?
On most real-world audio, yes. AssemblyAI's Universal models cut hallucination rate by roughly 30% versus Whisper and post an English mean WER of 5.6%, with the gap widening on noisy audio, accented speech, and proper nouns. Whisper holds up well on clean, single-speaker recordings.
Can you use both AssemblyAI and Whisper in the same application?
Yes, and many teams do—AssemblyAI for real-time features and difficult audio, Whisper for high-volume batch jobs, abstracted behind a common interface. AssemblyAI also offers a managed Whisper-Streaming option covering 99 languages, so you can keep Whisper's language coverage without hosting anything.
How long does it take to switch from Whisper to AssemblyAI?
Typically a few days—mostly API integration and deleting infrastructure code. Moving the other way requires GPU provisioning, model downloads, and rebuilding diarization, streaming, and redaction, which usually takes several weeks.
Which platform handles medical or legal terminology better?
AssemblyAI. Medical Mode activates with a single domain: "medical-v1" parameter, posts a 3.2% Missed Entity Rate on clinical entities, and costs $0.15/hr on top of the base model; keyterms prompting handles legal and other domain jargon with up to 1,000 supplied terms. Whisper generally needs fine-tuning for either domain. Note that Medical Mode is not available on the self-hosted deployment.
Lorem ipsum dolor sit amet, consectetur adipiscing elit, sed do eiusmod tempor incididunt ut labore et dolore magna aliqua. Ut enim ad minim veniam, quis nostrud exercitation ullamco laboris nisi ut aliquip ex ea commodo consequat. Duis aute irure dolor in reprehenderit in voluptate velit esse cillum dolore eu fugiat nulla pariatur.



