Insights & Use Cases
September 22, 2026

What is the best speech to text api to build ai medical ambient scribes?

Speech to text API for medical ambient scribes: compare real-time, HIPAA-compliant options with medical vocabulary, speaker diarization, and latency.

Kelsey Foster
, 
Growth
Reviewed by
No items found.
Table of contents

For ambient clinical capture, the API you want is the one with the lowest missed entity rate on medical terms, diarization measured by cpWER rather than by speaker count, real far-field handling, and a Business Associate Addendum you can actually sign. On those four criteria AssemblyAI’s Universal-3.5 Pro with Medical Mode posts a 3.2% Missed Entity Rate, averages cpWER 30.17 on diarization, exposes voice_focus: “far-field” for room audio, and offers a BAA that can be signed self-serve without a sales call.

The rest of this post explains why those are the four criteria, how the major APIs compare on them, and which AssemblyAI surface belongs in which part of a clinical workflow.

What an ambient scribe demands that dictation doesn’t

Dictation and ambient capture look like the same problem and are not. Dictation is one known speaker, close to the microphone, speaking deliberately into a device, producing a short utterance that is about to become text. Ambient capture is two or more speakers at unpredictable distances in an untreated room, talking over each other, with clinical meaning distributed across the whole conversation.

That difference generates four hard requirements that dictation never imposes.

1. Speaker attribution that survives crosstalk

A note built from an encounter is wrong if the patient’s symptom report is attributed to the clinician. The metric for this is cpWER (concatenated minimum-permutation word error rate), which scores transcription and speaker attribution together — so putting the right words on the wrong speaker costs you, which is exactly the failure mode that breaks a scribe.

Averaged across DiPCo, CALLHOME, NOTSOFAR and AMI, Universal-3.5 Pro posts cpWER 30.17, against Azure at 30.35, ElevenLabs Scribe V2 at 35.26, Speechmatics at 36.60, Gladia at 36.88 and Deepgram at 37.93. Azure is within 0.18 of us, so treat the published average as a shortlist rather than a verdict and run both on your own encounter audio. On the streaming path, diarization supports up to 10 speakers with revision — the model revises its speaker assignments as more audio arrives, with correction landing within roughly half a second of the stream ending. That matters for the family-member-in-the-room case, where a third voice appears thirty seconds in and a system that locked its speaker map early never recovers.

2. Far-field audio

A clinician wearing a headset and a microphone sitting on a desk two meters away are different acoustic problems. AssemblyAI exposes this as an explicit parameter rather than a generic noise setting: voice_focus takes “near-field” for headsets and phones, and “far-field” for rooms, kiosks, and similar open-mic environments. Note the hyphens — the underscored form is not a valid value. Strength is controlled by voice_focus_threshold, from 0.0 to 1.0, defaulting to 0.7 on the streaming surface. Voice Focus is a $0.10/hr add-on.

3. Medical entities, not words

Word error rate is the wrong scoreboard for clinical audio because it weights “the” and “metoprolol” identically. The metric that predicts whether a note is safe to sign is Missed Entity Rate — how often a clinically significant entity fails to appear at all. Medical Mode posts 3.2% MER, which is 87% fewer entity errors and roughly 20% fewer missed medical entities than the base model. Per-provider comparisons are on the benchmarks page; the metric itself is unpacked in MER vs. WER.

4. Context about the encounter

An ambient scribe knows things the model does not: the specialty, the reason for the visit, the patient’s active medication list. Feeding that in is a second accuracy lever that stacks with Medical Mode rather than competing with it.

On a public benchmark of 20,000 real voice-agent calls, detailed context cut medical-term entity errors by 43% and scenario-level context by 24% (prompting and keyterms). In internal testing, feeding a patient’s prior-visit note cut missed medical terms by a further 31%.

Which AssemblyAI surface for which clinical workflow

Because dictation and ambient capture are different problems, they route to different endpoints. This is the part most build decisions get wrong, so here it is explicitly.

Workflow Surface Price Medical Mode?
Live ambient capture during an encounter Universal-3.5 Pro Realtime $0.45/hr base, $0.60/hr with Medical Mode Yes
Post-encounter processing of recorded audio Universal-3.5 Pro (async) $0.21/hr base, $0.36/hr with Medical Mode Yes
Front-end short dictation into a note field Dictation API $0.62/hr flat, everything included No
Voice commands, IVR routing, push-to-talk Sync API $0.45/hr No

Ambient capture: Realtime, async, or both

Most production scribes run both paths. Streaming drives the live view the clinician watches during the encounter; async re-processes the full recording afterward for the note that gets signed, because a model that sees the whole file makes better diarization and entity decisions than one working turn by turn.

Running both with Medical Mode enabled costs $0.96 per hour of audio — $0.60/hr streaming plus $0.36/hr async. That is the number to put in a unit-economics model, and it is the whole transcription line.

import assemblyai as aai

aai.settings.api_key = "<YOUR_API_KEY>"

# Post-encounter pass: full recording, whole-file context
config = aai.TranscriptionConfig(
    speech_models=["universal-3-5-pro"],   # plural on async
    domain="medical-v1",
    speaker_labels=True,
)

transcript = aai.Transcriber().transcribe("encounter.wav", config)

On the streaming path the field is speech_model — singular — with the same domain: “medical-v1”, and voice_focus: “far-field” for room audio. Full parameters are in the streaming Medical Mode docs.

Front-end dictation: the Dictation API

When a clinician is dictating into a field — a chart note, an order, a message to a colleague — you do not want a verbatim transcript. Speakers restate and self-correct, and what was said is not what anyone wants to send. The Dictation API runs speech-to-text and an LLM cleanup pass in one call and returns both: text is the verbatim transcript and is never altered, llm_response is the cleaned version. It returns finished text in about 0.36 s, runs on Universal-3.5 Pro and accepts 32 language codes.

import assemblyai as aai

aai.settings.api_key = "<YOUR_API_KEY>"

config = aai.DictationConfig(
    stt_prompt="A doctor dictating a patient visit note.",
    keyterms_prompt=["amoxicillin", "lisinopril", "metoprolol"],
    llm_instruction=(
        "Remove filler words and rewrite as a concise clinical chart note."
    ),
)

result = aai.DictationTranscriber().transcribe_live("clip.wav", config)

transcript = result.text
rewrite = result.final_text  # cleaned text, falling back to the transcript

Two constraints to design around, stated plainly:

  • 120 seconds of audio per call. A routine SOAP note is about 250–400 words — two to three minutes of speech — so the cap covers short dictations and sits right at the edge of a typical note. Complex hospitalist or psychiatric notes run 800+ words and will not fit. Route those to pre-recorded STT instead.
  • No Medical Mode. domain: “medical-v1” runs on Universal-3.5 Pro async and Universal-3.5 Pro Realtime only. There is no domain parameter on the Dictation endpoint, and there is none on the Sync API either. If your accuracy story rests on the 3.2% MER figure, that claim belongs to the async and streaming paths — not to dictation. The Dictation API still runs on Universal-3.5 Pro and still takes keyterms_prompt for drug names, which covers a great deal of clinical vocabulary in practice, but it is a different capability.

One more thing worth knowing before you pick: clinicians are trained to speak punctuation — “Medications colon aspirin comma naproxen period.” AssemblyAI’s models apply punctuation automatically and handle spoken command words inconsistently, and there is no verbatim/spoken-form mode today. Deepgram ships a dictation=true flag that converts spoken commands into punctuation and markets it in their medical content. The documented workaround on our side is an LLM pass over each finalized turn — via the LLM Gateway — to interpret and strip command words, which is the pattern in the medical scribe best practices guide. It works well, and it is a workaround rather than a native mode.

Start Free On Clinical Audio

The free tier covers 185 hours of pre-recorded and 333 hours of streaming transcription — enough to run a real ambient pilot rather than a demo.

Sign up free

How the major speech-to-text APIs compare for ambient capture

Below is what we can substantiate. AssemblyAI figures come from published benchmarks. Where we do not have a sourced, current figure for another vendor — and particularly where a vendor’s compliance posture is concerned — the cell reads “Not publicly stated” rather than a guess. Check each vendor’s own documentation before you commit; we are not going to characterize someone else’s contract terms.

Criterion AssemblyAI (Universal-3.5 Pro / Realtime) Deepgram AWS Transcribe Medical Google Azure Whisper
Medical entity accuracy (Missed Entity Rate) 3.2% MER with Medical Mode Included in our benchmark set — see benchmarks Included in our benchmark set — see benchmarks Included in our benchmark set — see benchmarks Included in our benchmark set — see benchmarks Not benchmarked by AssemblyAI; not publicly stated
Entity error rate, live conversational audio (Pipecat open STT benchmark, lower is better — general entities: names, places, phone numbers, not clinical terms) 15.31% 50.50% (Flux) Not in this benchmark 21.51% (Chirp3) Not in this benchmark Not in this benchmark
Word error rate, live conversational audio (same benchmark) 6.99% 15.58% (Flux) Not in this benchmark 9.04% (Chirp3) Not in this benchmark Not in this benchmark
Diarization quality (cpWER, lower is better) 30.17; up to 10 speakers on streaming, with revision within ~0.5 s of stream end 37.93 Not publicly stated Not publicly stated 30.35 Not publicly stated
Far-field handling Explicit voice_focus: “near-field” / “far-field” parameter (+$0.10/hr) Not publicly stated Not publicly stated Not publicly stated Not publicly stated Not publicly stated
Streaming support Universal-3.5 Pro Realtime, $0.45/hr base Streaming model (Flux) benchmarked above; pricing not publicly stated here Not publicly stated Streaming model (Chirp3) benchmarked above; pricing not publicly stated here Not publicly stated Not publicly stated
BAA availability BAA available; signable self-serve without a sales call Not publicly stated — confirm with the vendor Not publicly stated — confirm with the vendor Not publicly stated — confirm with the vendor Not publicly stated — confirm with the vendor Not publicly stated — confirm with the vendor
Spoken-punctuation / verbatim dictation mode No native mode today; documented LLM-pass workaround Ships a dictation=true flag that converts spoken commands into punctuation Not publicly stated Not publicly stated Not publicly stated Not publicly stated

Three things to read out of that table.

First, the gap between providers is much wider on entity errors than on word errors. On the Pipecat benchmark, AssemblyAI’s word error rate is roughly half the next-best figure but its entity error rate is a third of it. Entities are where the clinically load-bearing information lives, which is why a headline WER number can make two APIs look closer than they are for this use case. Note that the Pipecat entity figures cover general conversational entities — names, places, phone numbers — rather than clinical terms.

Second, diarization is the row where the field is closest. Azure is 0.18 cpWER behind us. If speaker attribution is the requirement that decides your build, that is a genuine two-horse race and you should run both on your own exam-room audio.

Third, we listed Deepgram’s dictation=true flag as a genuine advantage because it is one. A comparison table that hides a competitor’s real capability is worthless to the person reading it, and this buyer will find out within a week of evaluating.

What teams building scribes actually run into

Commure builds clinical workflow software including an ambient product.

“We’ve integrated the newest models from AssemblyAI for pre-recorded audio ASR in our ambient product, and it’s been excellent…”

— Gautam Pradeep, Tech Lead, Commure

Other teams building in this space on AssemblyAI include Sully AI, Heidi Health, DeepScribe, Knowtex, Magentus Healthcare, Chapter, and NMDP. The patterns that recur across them:

  • The room, not the model, is the hard part. Exam rooms have hard surfaces, HVAC noise, and a microphone nobody positions correctly. This is what voice_focus: “far-field” is for, and it is the first parameter to change when accuracy looks worse in production than in testing.
  • Diarization errors surface as clinical errors. A misattributed symptom report is a worse failure than a misspelled word, and it will not be caught by a WER-based eval. Evaluate with cpWER.
  • Context is free accuracy. If your product already knows the specialty and the medication list, passing them in is the cheapest accuracy improvement available — 43% fewer medical-term entity errors on the public benchmark.
  • The signed note is not the live note. Re-processing the recording asynchronously after the encounter produces a better document than the streaming transcript alone. That is what the $0.96/hr combined figure buys.
Test It On Your Own Audio

Run Medical Mode against a real encounter recording without any integration work, and hear what far-field capture and diarization do to your own exam-room audio.

Try playground

PHI, BAAs, and what you can and cannot say

AssemblyAI is a business associate under HIPAA, not a covered entity. The legal-approved description:

AssemblyAI enables covered entities and their business associates subject to HIPAA to use the AssemblyAI services to process protected health information (PHI). AssemblyAI is considered a business associate under HIPAA, and we offer a standard Business Associate Addendum (BAA) that is required under HIPAA to ensure that AssemblyAI appropriately safeguards PHI.

The BAA can be reviewed and signed self-serve from the Data Controls page in the dashboard, without a sales call — see can you sign a BAA and the Business Associate Addendum. Supporting controls include PHI redaction across audio and transcripts, SOC 2 Type 2, ISO 27001:2022, and PCI DSS v4.0. Deployment options cover hosted, self-hosted in your own cloud account, and EU data residency at the same price.

Talk Through A Clinical Deployment

Walk through the BAA, EU data residency and self-hosted options with someone who has shipped ambient scribes before. The end-to-end medical scribe example is the fastest way to see the build first.

Talk to AI expert

Frequently asked questions

Is AssemblyAI HIPAA-compliant?

AssemblyAI offers a standard Business Associate Addendum (BAA), which is what HIPAA requires of a vendor processing PHI on a covered entity’s behalf — AssemblyAI is a business associate under HIPAA, not a covered entity. The BAA can be reviewed and signed self-serve without a sales call. Supporting controls include PHI redaction across audio and transcripts, SOC 2 Type 2, ISO 27001:2022, and PCI DSS v4.0.

What is the best speech-to-text API for an AI medical scribe?

The one with the lowest missed entity rate on clinical terms and diarization measured by cpWER. AssemblyAI’s Universal-3.5 Pro with Medical Mode posts a 3.2% Missed Entity Rate and averages cpWER 30.17 on speaker attribution, with Azure the closest competitor on diarization at 30.35. Word error rate alone is a poor proxy, because it weights common words and drug names identically. Run your own encounter audio through the shortlist before deciding.

How much does it cost to run an ambient scribe?

$0.96 per hour of audio if you run both paths with Medical Mode: $0.60/hr for streaming (Universal-3.5 Pro Realtime at $0.45/hr plus the $0.15/hr Medical Mode add-on) and $0.36/hr for the async re-processing pass (Universal-3.5 Pro at $0.21/hr plus $0.15/hr). Async-only is $0.36/hr. Optional add-ons such as diarization and Voice Focus are itemized on the pricing page.

How do you handle multiple speakers in an exam room?

Streaming diarization supports up to 10 speakers and revises its assignments as more audio arrives, with correction landing within roughly half a second of the stream ending. That revision step is what handles the common case of a family member joining a conversation partway through. Diarization quality is best evaluated with cpWER, which scores transcription and speaker attribution together; Universal-3.5 Pro averages 30.17, ahead of ElevenLabs Scribe V2 at 35.26 and Deepgram at 37.93, with Azure closest at 30.35.

Does the model handle a microphone across the room?

Yes — set voice_focus: “far-field”, which is tuned for rooms and open-mic environments, as opposed to “near-field” for headsets and phones. Note the hyphens; the underscored form is not a valid value. Strength is controlled by voice_focus_threshold from 0.0 to 1.0, defaulting to 0.7 on streaming. Voice Focus is a $0.10/hr add-on, and it is the first parameter to change if production accuracy in exam rooms is worse than your bench testing predicted.

Can I improve accuracy on a specific patient or specialty?

Yes, and it stacks with Medical Mode rather than replacing it. Passing encounter context — specialty, visit reason, active medication list — cut medical-term entity errors by 43% on a public benchmark of 20,000 real calls, with scenario-level context alone cutting them by 24%. In internal testing, feeding a patient’s prior-visit note cut missed medical terms by a further 31%. See the prompting and keyterms documentation.

Title goes here

Lorem ipsum dolor sit amet, consectetur adipiscing elit, sed do eiusmod tempor incididunt ut labore et dolore magna aliqua. Ut enim ad minim veniam, quis nostrud exercitation ullamco laboris nisi ut aliquip ex ea commodo consequat. Duis aute irure dolor in reprehenderit in voluptate velit esse cillum dolore eu fugiat nulla pariatur.

Button Text
Medical
ambient AI scribe