Run the arithmetic before the pilot. A clinic doing 20 visits a day at 20 minutes each produces about 6.7 hours of encounter audio. Transcribed with Universal-3.5 Pro and Medical Mode at $0.36/hr, that's roughly $2.40 a day. Per clinician. The transcription layer of an AI documentation program is, in almost every deployment we see, the cheapest line item in it.
Which means the interesting questions in healthcare AI transcription are not about cost. They're about accuracy at the entity level, speaker attribution in a room with four people in it, PHI governance, and whether clinicians trust the output enough to stop rewriting it. Get those right and the economics take care of themselves. Get them wrong and you've built an expensive proofreading job.
This is a practical overview: how the technology actually works, where it's being deployed, what it costs to run, how to think about clinical safety, and what to put in a vendor evaluation. Written from the API side of the problem, since that's where we sit.
How AI medical transcription works
Four stages, and each one can fail independently.
1. Speech recognition
Audio in, text out. This is where clinical vocabulary either survives or doesn't. General models resolve acoustic ambiguity toward common everyday words, which is exactly wrong when the intended word is "hydralazine." A clinical mode shifts those priors — Medical Mode on Universal-3.5 Pro records a 3.2% Missed Entity Rate, roughly 20% fewer missed medical entities than the base model and 87% fewer entity errors.
2. Speaker separation
A clinical transcript without attribution is barely useful — you can't tell a patient's symptom report from a clinician's instruction. Universal-3.5 Pro's diarization is optimized for cpWER rather than diarization error rate, which means it's scored on whether the right words land under the right speaker, and it holds up on short turns and overlapped speech. Streaming supports up to 10 speakers with revision as more audio arrives.
3. Clinical understanding
Turning a transcript into structured findings — problems, medications, plan — is an LLM task layered on top. Worth being clear-eyed about the dependency: this stage cannot detect or repair an error from stage one. A summarizer handed the wrong drug name will write the wrong drug name into the assessment with total confidence.
4. EHR integration
The last mile, and usually the longest. This is where most timelines slip, and it has nothing to do with the speech model.
Where it's actually being deployed
Ambient documentation. A microphone captures the encounter and a note comes out. The highest-value and hardest deployment — far-field audio, multiple speakers, unscripted patient language. Our guide to speech-to-text for ambient scribes covers the technical requirements.
Dictation. Clinician-driven, cooperative audio, immediate feedback. The easiest to deploy and the easiest to get adoption on, because the clinician stays in control.
Post-visit processing. Batch transcription of recorded encounters, procedure reports, and correspondence. No latency constraint means you get maximum model quality at the lowest rate.
Care coordination and telehealth. Multi-party calls, case conferences, care team handoffs — where diarization quality matters more than anywhere else because attribution is the whole point.
Patient-facing phone lines. Scheduling, refills, intake, triage routing. This is a voice agent problem rather than a transcription problem, and it's covered below.
What it costs to run
Straight arithmetic from list rates. Substitute your own volumes.
| Configuration | Rate per audio hour | Typical use |
|---|---|---|
| Universal-3.5 Pro, async | $0.21 | General recorded audio, no clinical entities |
| Universal-3.5 Pro + Medical Mode, async | $0.36 | Post-visit notes, ambient scribe batch pass |
| Universal-3.6 Pro Realtime | $0.45 | Live captions, general streaming |
| Universal-3.6 Pro Realtime + Medical Mode | $0.60 | Live clinical documentation, in-visit prompts |
| Voice Agent API | $4.50 flat | Patient phone lines, scheduling, intake |
Current rates always live on pricing. Two things worth noticing.
First, running both a live stream during the visit and a full async pass afterward costs $0.96/hr of encounter audio. That's a pattern worth considering: the clinician gets live text, and the note that actually gets signed benefits from full-document context and revised diarization. Doubling the transcription bill is not a meaningful budget event at these rates.
Second, the real costs in a documentation program are elsewhere — EHR integration engineering, clinician training, template design, quality review, change management. Budget accordingly, and don't let a procurement process spend three months negotiating a line item that rounds to nothing.
Sign up free, transcribe a day of real encounters with Medical Mode on, and read the actual cost and accuracy off your own data.
Accuracy and clinical safety
The single most common measurement mistake in healthcare AI procurement is running a word error rate bakeoff. WER treats every token as equal weight, so "the" and "amiodarone" count the same. A model can win a WER comparison while being systematically worse at the fifteen words per encounter that carry clinical meaning.
Missed Entity Rate scores only those words — drugs, dosages, conditions, procedures, anatomy. On our benchmarks, Universal-3.5 Pro with Medical Mode records the lowest MER across benchmarked providers at 3.2%.
The error classes to build guardrails around
- Sound-alike drugs. Hydralazine and hydroxyzine, clonidine and Klonopin, metoprolol and metronidazole. One or two phonemes apart, and those are the phonemes that vanish under room noise.
- Negation loss. "No history of MI" dropping the "no." Fluent, plausible, and meaning-inverting — the class that worries clinical safety reviewers most.
- Dosage and number handling. Where the model's formatting choices become clinical assertions.
- Attribution flips. A caregiver's speculation recorded as patient history.
Design review workflows around these four specifically rather than asking clinicians to proofread everything equally. Attention is the scarce resource, and pointing it at the high-risk categories is worth more than a blanket sign-off.
Contextual prompting is the underused lever
You can pass free-text context to the model before it transcribes. In a clinical setting, the highest-value context is the patient's own prior-visit note — and in an internal healthcare test, feeding exactly that cut missed medical terms by 31%.
This is the accuracy improvement that costs nothing. The data is already in your system. If you're specifying an AI documentation program right now, make context assembly a requirement in the architecture rather than a phase-two nice-to-have; retrofitting it later means touching every integration point.
Implementation
The clinical configuration is a parameter, not a project:
import assemblyai as aai
aai.settings.api_key = "YOUR_API_KEY"
config = aai.TranscriptionConfig(
speech_models=["universal-3-5-pro"],
domain="medical-v1",
speaker_labels=True,
speakers_expected=3,
)
transcript = aai.Transcriber().transcribe("encounter.wav", config=config)
for utterance in transcript.utterances:
print(f"Speaker {utterance.speaker}: {utterance.text}")Streaming uses speech_model singular against wss://streaming.assemblyai.com/v3/ws:
from assemblyai.streaming.v3 import (
StreamingClient,
StreamingClientOptions,
StreamingParameters,
)
client = StreamingClient(
StreamingClientOptions(api_key="YOUR_API_KEY")
)
client.connect(
StreamingParameters(
sample_rate=16000, speech_model="universal-3-6-pro",
domain="medical-v1",
voice_focus="far-field",
mode="balanced",
speaker_labels=True,
)
)Universal-3.6 Pro Realtime defaults turn detection to min_turn_silence 128ms and max_turn_silence 1280ms on the balanced preset, and gives you three modes — min_latency, balanced, max_accuracy — so you can tune per surface. Full reference in the docs.
Rollout advice that has nothing to do with the model
Start with one specialty and one workflow. Pick a service line with high documentation burden and cooperative clinicians, and let them be the reference. Measure baseline documentation time before you deploy anything, because without a baseline you'll be arguing about whether it helped for a year. Give clinicians a visible correction path and actually feed those corrections into your context assembly. And decide in advance what "good enough" looks like numerically, or the pilot never ends.
Toggle Medical Mode, far-field voice focus, and diarization on the same encounter in the playground before you spec the pipeline.
PHI, BAA, and governance
Compliance is a gate, not a differentiator — but failing the gate ends the project.
AssemblyAI signs a Business Associate Addendum (BAA) for customers processing PHI and acts as a business associate under HIPAA for that data. The process is documented in our BAA FAQ and the agreement itself is at legal/business-associate-agreement. We're SOC 2 Type 2.
Operationally, the pieces to specify: PHI redaction runs across both audio and transcripts, which matters because a redacted transcript sitting next to an unredacted recording isn't a redacted encounter. EU data residency is available at api.eu.assemblyai.com. Self-hosted deployment is available when audio can't leave your environment at all.
The governance questions to answer internally, before procurement asks: who can access recordings and for how long, what the patient consent flow looks like, how corrections are audited, and what happens to the audio after the note is signed. These are policy decisions, not vendor features, and they take longer to settle than the technical integration.
Vendor evaluation criteria
What we'd put on the scorecard, weighted roughly in this order:
- Missed Entity Rate on your own audio. Not the vendor's benchmark. Yours.
- Diarization on short turns. Score attribution on utterances under five words specifically.
- Context handling. Free-text context or a keyword list you have to maintain forever? This is an operational cost question disguised as a feature question.
- Redaction scope, retention, and training use. In writing.
- Deployment options. Hosted, EU residency, self-hosted.
- Contract structure. Commit size and overage terms matter more than list rate.
For the provider-by-provider view, see our comparison of the best medical speech-to-text options and the head-to-head on AssemblyAI vs Deepgram for medical transcription. The broader API selection process is in our medical transcription API guide.
Organizations building clinical voice products on this stack include Commure, Sully AI, Heidi Health, Deepscribe, Knowtex, Magentus Healthcare, Chapter, and NMDP — ambient documentation, clinical workflow, Medicare navigation, donor matching. Different products, one shared requirement.
Beyond documentation: the patient phone line
Documentation gets the attention, but the higher-volume voice problem in most health systems is the phone. Scheduling, refill requests, prior authorization follow-up, intake, triage routing — all of it currently absorbing staff time.
The Voice Agent API handles this as a single WebSocket at a flat $4.50/hr, replacing separate speech-to-text, LLM, and text-to-speech vendors and the orchestration glue between them. For healthcare the relevant detail is that the entity accuracy that matters on a refill call — distinguishing amoxicillin from amoxicillin-clavulanate — comes from the same model family as your clinical documentation, rather than from a general-purpose model that has never seen a formulary.
What changes over the next two years
The accuracy race is flattening. The top of the medical benchmark is crowded, and the difference between leading models on clean audio is narrowing to the point where it won't be the deciding factor in a purchase.
What's opening up instead is a gap between organizations based on how well they feed their models. Contextual prompting means transcription accuracy is now partly a function of your own data infrastructure — whether the prior-visit note, the medication list, and the specialty context reach the model at request time. A 31% reduction in missed medical terms from a single note is the shallow end of that curve. The health systems and vendors that treat context assembly as core architecture will see their accuracy improve without changing vendors at all, while everyone else keeps re-running bakeoffs chasing a percentage point they could have gotten from their own database. That's a build decision you're making right now, whether or not you're making it deliberately. See medical solutions for how the pieces fit.
Bring your encounter volumes, workflow constraints, and governance requirements. We’ll help you model the cost, design the pilot, and start the BAA in parallel.
Frequently asked questions
How accurate is AI medical transcription?
Measured properly — on Missed Entity Rate rather than word error rate — Universal-3.5 Pro with Medical Mode records 3.2% MER, the lowest across benchmarked providers. That's also 87% fewer entity errors than the same model without Medical Mode. See benchmarks for methodology.
How much does AI medical transcription cost?
The transcription layer is $0.36/hr of audio for async with Medical Mode ($0.21 base plus $0.15) and $0.60/hr for streaming with Medical Mode. A clinician seeing 20 patients for 20 minutes each generates about 6.7 hours of audio, so roughly $2.40 a day. The larger costs in a documentation program are integration, training, and change management — not the API.
How does AssemblyAI handle HIPAA and PHI?
AssemblyAI signs a Business Associate Addendum (BAA) for customers processing PHI and operates as a business associate under HIPAA for that data. We're SOC 2 Type 2, PHI redaction covers both audio and transcripts, and EU data residency and self-hosted deployment are available. Start with our BAA FAQ.
How does AssemblyAI secure patient data?
Four layers: a signed BAA making us a business associate for your PHI, SOC 2 Type 2 controls, PHI redaction across both audio and transcript, and deployment options including EU data residency at api.eu.assemblyai.com and self-hosting for audio that can't leave your infrastructure. Ask us in writing about retention periods and whether your audio is used for model training — we're direct about both.
Does it work across specialties and languages?
Medical Mode's clinical entity layer covers English, Spanish, German, and French, for both pre-recorded and streaming audio. Separately, the base model code-switches natively across 18 languages with no configuration, so a patient alternating between languages mid-sentence doesn't break the transcript. Specialty coverage spans drugs, conditions, and procedures broadly rather than being tuned per specialty — and contextual prompting is how you add specialty-specific vocabulary without maintaining a list.
What happens when the AI makes a transcription error?
You design for it. The four error classes worth building explicit review around are sound-alike drug names, lost negations, dosage and number formatting, and speaker attribution flips. Point clinician attention at those categories rather than asking for a blanket proofread, feed corrections back into your context assembly, and remember that a downstream LLM cannot detect an upstream entity error — it will propagate it confidently. See medical transcription use cases.