Insights & Use Cases
September 22, 2026

Best medical speech recognition software and APIs in 2026

Medical speech recognition software turns clinical speech into accurate, structured notes. Compare the 8 best tools and APIs for 2026 on medical accuracy, streaming, BAA availability, and pricing.

Kelsey Foster
, 
Growth
Reviewed by
No items found.
Table of contents

Medical speech recognition is evaluated on one number that general transcription benchmarks do not report: how often the system drops a clinically meaningful term. A transcript can post an excellent word error rate and still lose the drug name, the dosage, or the laterality — and those are the only errors that change patient care.

That number is Missed Entity Rate (MER). AssemblyAI’s Medical Mode posts a 3.2% MER. This guide compares eight options on MER and on the four other dimensions that actually separate them: latency, deployment, language coverage, and whether the vendor will sign a Business Associate Addendum.

It is written for the person choosing infrastructure — a developer, an engineering lead, a technical buyer at a healthcare software company. If you are a clinician or a practice shopping for an app to dictate into, start with our Dragon Medical alternatives guide instead, which covers the per-clinician software market directly.

What the AI Overview gets wrong about medical dictation software

Search “medical dictation software” today and Google’s AI Overview returns a list that mixes two products that are not substitutes for each other. It will name Dragon Medical One next to Amazon Transcribe Medical, quote a monthly per-clinician price next to a per-hour usage rate, and present them as competing options on a single list.

They are not competing options. They are two different purchases with different buyers, different price units, and different work left over for you to do.

Path 1: clinician-facing applications

These are finished products a physician opens and talks into. Dragon Medical One, Suki, Abridge, DeepScribe and Augmedix live here.

  • Price unit: per clinician, per month — typically in the $49–99 range.
  • What’s included: the desktop or mobile client, the microphone workflow, EHR integration, training, and a support desk.
  • Who buys: practices, health systems, and individual clinicians.
  • What you build: nothing.

Path 2: developer speech APIs

These are endpoints. AssemblyAI, Deepgram, Speechmatics, Amazon Transcribe Medical, Google, NVIDIA Riva, Rev AI and ElevenLabs live here.

  • Price unit: per hour of audio processed — typically $0.15 to $0.62 per hour, with clinical tuning as an add-on.
  • What’s included: transcription, and the parameters that shape it.
  • Who buys: healthcare software companies, AI scribe startups, EHR vendors, telehealth platforms.
  • What you build: the application — audio capture, the interface, error handling, the EHR write-back, and the clinical QA around all of it.

The two prices are not comparable because they are not buying the same thing. A clinician dictating 45 minutes a day generates roughly 15 hours of audio a month, which costs under $10 at API rates and $79 or more as a subscription. The difference is not margin; it is the entire application layer. The right question is not “which is cheaper” but “am I the end user, or am I shipping this to end users?” Everything else follows from that answer.

The rest of this guide covers Path 2, because that is where the choice is genuinely technical and genuinely hard to reverse.

How to evaluate medical speech recognition: five criteria

1. Missed Entity Rate, not word error rate

WER counts every word equally. In a clinical transcript they are not equal: “the” and “metoprolol” carry very different consequences. MER isolates the medical entities — drugs, dosages, conditions, procedures, anatomy — and measures how often they are missed. Two systems can post near-identical WER and differ by a factor of several on MER. We walk through the divergence in detail in MER vs. WER.

2. Whether there is a clinical mode at all

Some providers ship a medical-tuned model or domain flag; others expect you to compensate with custom vocabulary lists. A tuned model is a materially different thing from a word-boost list, and the gap shows up on exactly the terms that matter. On AssemblyAI, Medical Mode is one parameter — domain: “medical-v1” — and it delivers 87% fewer entity errors than the base model.

3. Latency, matched to the workflow

Ambient scribing and post-visit documentation tolerate batch processing. Dictation does not: the clinician is waiting for text to appear, and anything over a second reads as broken. Ask for p50, not average, and ask what it is measured against — request-to-finished-text, or first partial.

4. Deployment and data residency

Whether audio can leave your environment is often decided before the evaluation starts. Options range from hosted cloud, to regional data zones, to fully self-hosted inside your own VPC. Confirm which the vendor actually offers on the plan you will be on, not on the enterprise tier you may never buy.

5. The BAA, in writing, before the pilot

HIPAA does not certify software. It places obligations on covered entities and on their business associates, and the contractual instrument that binds a vendor is the Business Associate Addendum. AssemblyAI enables covered entities and their business associates subject to HIPAA to use the AssemblyAI services to process protected health information (PHI). AssemblyAI is considered a business associate under HIPAA, and we offer a standard Business Associate Addendum (BAA) that is required under HIPAA to ensure that AssemblyAI appropriately safeguards PHI — and it can be signed self-serve without a sales call.

For every other vendor, ask them directly. We do not characterize another company’s contractual posture in the table below, and you should not accept anyone else’s characterization of it either.

The 8 tools compared

Provider What it is Clinical tuning Pricing model Deployment Dictation-specific API BAA offered
AssemblyAI Voice AI infrastructure — pre-recorded, streaming, sync and dictation APIs Medical Mode (domain: "medical-v1"), 3.2% MER $0.21/hr async, $0.45/hr streaming; Medical Mode +$0.15/hr Hosted cloud, EU data residency, self-hosted Yes — $0.62/hr Yes
Deepgram Speech-to-text API with real-time and batch models Medical models marketed; included in our published MER benchmark Usage-based Cloud and self-hosted dictation=true flag that converts spoken commands to punctuation Not publicly stated
Speechmatics Speech-to-text API with broad language coverage Included in our published MER benchmark Usage-based Cloud and on-premises Not offered Not publicly stated
Amazon Transcribe Medical AWS’s medical speech-to-text service Purpose-built medical service; included in our published MER benchmark Usage-based, per AWS’s published rates AWS regions Not offered Not publicly stated
Google Cloud Speech-to-Text (Chirp 3) Google’s speech recognition API Included in our published MER benchmark; general-purpose model Usage-based Google Cloud regions Not offered Not publicly stated
NVIDIA Riva Self-hosted speech SDK you run on your own GPUs Not publicly stated Infrastructure cost — you run the hardware Self-hosted / on-premises only Not offered Not applicable — you host it
Rev AI Speech-to-text API from a human-transcription business Not publicly stated Usage-based Cloud Not offered Not publicly stated
ElevenLabs Scribe v2 Speech-to-text model from a voice-AI company Not publicly stated Usage-based Cloud Not offered Not publicly stated

Two notes on reading this table. “Not publicly stated” means exactly that — we did not find a published claim we could cite, and we are not going to infer one. And clinical tuning entries referencing our benchmark mean the provider is included in the head-to-head medical entity comparison published at assemblyai.com/benchmarks, with the audio sets and methodology alongside the numbers.

1. AssemblyAI

AssemblyAI is Voice AI infrastructure with four transcription surfaces — pre-recorded, streaming, sync, and dictation — sharing one flagship model family and one clinical tuning mode.

Medical Mode

Medical Mode is a single parameter that tunes the model for clinical vocabulary. It runs on Universal-3.5 Pro for pre-recorded audio and Universal-3.5 Pro Realtime for streaming, in English, Spanish, German and French.

  • 3.2% Missed Entity Rate (absolute)
  • 87% fewer entity errors than the base model
  • $0.15/hr add-on: $0.36/hr combined with pre-recorded, $0.60/hr combined with streaming
import assemblyai as aai

aai.settings.api_key = "<YOUR_API_KEY>"

config = aai.TranscriptionConfig(
    speech_models=["universal-3-5-pro"],
    domain="medical-v1",
)

transcript = aai.Transcriber().transcribe("visit-note.wav", config)
print(transcript.text)

Streaming takes speech_model (singular) with the same domain. Parameters are documented for pre-recorded and streaming separately.

Medical Mode pairs with contextual prompting, and the two compound rather than overlap: Medical Mode tunes the model for clinical language generally, while prompting tells it about this specific encounter. On a public benchmark of 20,000 real voice-agent calls, detailed context cut medical-term entity errors by 43% and scenario-level context by 24% (documented here).

The Dictation API for short-form clinical dictation

The Dictation API is a separate product for a separate job. Transcription models are verbatim by design, but what was said is not what anyone wants to send. Dictation pairs speech-to-text with an LLM cleanup pass in one low-latency call: self-corrections resolve to what the speaker landed on, filler disappears, names come back spelled correctly, and the verbatim transcript is returned alongside the cleaned version so nothing is lost.

  • $0.62/hr flat — every feature included, one line on the bill
  • 0.36 s to finished text; typical short clips return in under a second, and every response carries request_time_ms
  • Runs on Universal-3.5 Pro, across 32 languages
  • Three levers: stt_prompt for the setting, keyterms_prompt for exact terms, llm_instruction for the rewrite task
import assemblyai as aai

aai.settings.api_key = "<YOUR_API_KEY>"

config = aai.DictationConfig(
    stt_prompt="A doctor dictating a patient visit note.",
    keyterms_prompt=["amoxicillin", "lisinopril", "metoprolol"],
    llm_instruction=(
        "Remove filler words and rewrite as a concise clinical chart note."
    ),
)

result = aai.DictationTranscriber().transcribe_live("clip.wav", config)

transcript = result.text
rewrite = result.final_text  # the cleaned-up text, falling back to the transcript

Where it stops, stated plainly. Audio is capped at 120 seconds per call. A routine SOAP note runs roughly 250–400 words — two to three minutes at 150 wpm — so 120 seconds covers short dictations and sits at the edge of a typical note; complex hospitalist and psychiatric notes run 800+ words and will not fit. And Medical Mode is not available on the Dictation API: there is no domain parameter, and the same applies to the Sync API. If you are held to a missed-entity number, that figure lives on the pre-recorded and streaming paths. Route longer or entity-critical dictation to Pre-recorded STT or real-time streaming with Medical Mode, and use the Dictation API for the short-form case it was built for.

Spoken punctuation

One more limit worth naming, because clinicians ask about it immediately. Many physicians are trained to speak punctuation — “Medications colon aspirin comma naproxen period.” AssemblyAI’s models apply punctuation automatically and handle spoken command words inconsistently; there is no verbatim or spoken-form mode today. Deepgram ships a dictation=true flag that does convert spoken commands into punctuation, and if that is a hard requirement for your users it is a genuine advantage.

The documented workaround on our side is to run an LLM pass over each finalized turn through the LLM Gateway to interpret and strip the command words, which is the pattern in our guide to building a medical scribe. It works, and it is fast: Qwen3.5 4B Fast (qwen3.5-4b-32k-fast) is the one model on the Gateway roster that AssemblyAI hosts on its own GPUs, averaging 612 ms on voice rewrite tasks at $0.10/$0.50 per million prompt/completion tokens. It is still a workaround rather than a native mode, and we would rather you know that before the pilot than after.

In production

“We’ve integrated the newest models from AssemblyAI for pre-recorded audio ASR in our ambient product, and it’s been excellent…”

— Gautam Pradeep, Tech Lead, Commure

Other healthcare teams building on the platform include Sully AI, Heidi Health, Knowtex, Magentus Healthcare, NMDP and Chapter.

Test It On Your Own Clinical Audio

Get 185 hours of pre-recorded transcription and 333 hours of streaming on the free tier — enough for a real clinical pilot, not a demo.

Sign up free

2. Deepgram

A speech-to-text API with both batch and real-time models, marketed into healthcare and included in our published MER benchmark.

The genuine differentiator: the dictation=true flag, which converts spoken command words into punctuation. If your clinicians dictate punctuation out loud and you cannot add a post-processing step, this is the one capability on this list that no amount of prompt engineering fully replaces. We say so in our own comparison because a table that hides it is useless to the person reading it.

Trade-off to weigh: on the Pipecat open STT benchmark run against real agent conversations, Deepgram Flux posts a 15.58% word error rate and a 50.50% entity error rate against Universal-3.5 Pro Realtime’s 6.99% and 15.31%. For a deeper head-to-head see AssemblyAI vs. Deepgram for medical transcription.

3. Speechmatics

A speech-to-text API known for broad language coverage, available in cloud and on-premises deployments, and included in our published MER comparison — where it is currently the strongest competitor on medical entity accuracy specifically.

Best for: teams whose primary constraint is language breadth or an on-premises requirement, and teams for whom medical entity accuracy is the single scored requirement.

Watch for: no dictation-specific endpoint. Benchmark it on your own medical audio alongside AssemblyAI rather than assuming either general accuracy or our benchmark transfers to your case.

4. Amazon Transcribe Medical

AWS’s purpose-built medical speech-to-text service — one of the few options on this list that ships an explicitly medical product rather than a general model you tune.

Best for: teams already standardized on AWS who want medical transcription billed on the same account, inside the same VPC, under the same agreement. That procurement simplicity is a real advantage and frequently decides the evaluation.

Watch for: it is included in our MER benchmark. If entity accuracy is your scored requirement, run both on your own audio rather than deciding on the AWS-native convenience alone.

5. Google Cloud Speech-to-Text (Chirp 3)

Google’s general-purpose speech recognition API, included in our MER benchmark and in the Pipecat voice-agent benchmark, where Chirp 3 posts a 9.04% word error rate and a 21.51% entity error rate against Universal-3.5 Pro Realtime’s 6.99% and 15.31%.

Best for: teams already on Google Cloud, or workloads that need very wide language coverage on general-domain audio.

Watch for: it is a general model. There is no clinical domain flag equivalent to Medical Mode, so clinical vocabulary handling depends on what you can achieve with custom vocabulary and context.

6. NVIDIA Riva

Not a hosted API — a speech SDK you deploy and run on your own GPUs. The pricing model is your infrastructure bill.

Best for: organizations with a hard requirement that audio never leaves their own hardware, an existing GPU fleet, and the MLOps capacity to operate speech models in production.

Watch for: the total cost is engineering time and GPU capacity, not a per-hour rate, and it is easy to underestimate both. If self-hosting is the requirement rather than the preference, note that AssemblyAI also offers a self-hosted deployment — for example inside a customer’s own AWS VPC — so you can get data-residency control without owning the model operations.

7. Rev AI

A speech-to-text API from a company whose core business is human transcription, which shapes the product: the API is one path, and human-in-the-loop review is another.

Best for: workflows where some transcripts need a human pass and you would rather buy both from one vendor.

Watch for: no publicly stated clinical tuning mode and no dictation-specific endpoint. Validate medical entity handling directly.

8. ElevenLabs Scribe v2

A speech-to-text model from a company better known for voice synthesis, and a reasonable general-purpose transcription option.

Benchmark context: on the Pipecat voice-agent benchmark, Scribe v2 posts a 9.76% word error rate and a 39.70% entity error rate, against Universal-3.5 Pro Realtime’s 6.99% and 15.31%. On the async code-switching benchmark published in the Universal-3.5 Pro launch post it averages 8.77 normalized WER across five language pairs, against Universal-3.5 Pro’s 7.69.

Watch for: no publicly stated clinical mode. The entity error gap is the number to pay attention to for clinical work, since entities are what clinical transcripts are made of.

Compare Output Side By Side

Run a dictated chart note through the playground and toggle Medical Mode on and off before you commit to an integration.

Try playground

Picking one

  • Entity accuracy is the scored requirement. Benchmark Medical Mode on Universal-3.5 Pro against Speechmatics and Amazon Transcribe Medical on your own audio. Published benchmarks narrow the field; they do not settle it.
  • Clinicians dictate punctuation out loud. Score that capability explicitly. Deepgram’s dictation=true is native; on AssemblyAI it is an LLM post-processing pass.
  • Dictations are short and go straight into a text field. The Dictation API at $0.62/hr with 0.36 s to finished text is the purpose-built path — within the 120-second cap and without Medical Mode.
  • Audio cannot leave your environment. NVIDIA Riva if you want to own the stack; AssemblyAI self-hosted or EU data residency if you want the control without the model operations.
  • You are on AWS or Google Cloud and procurement is the constraint. The native option will usually win on friction. Just run the MER comparison first so you know what you are trading.

If you are evaluating against Dragon specifically, or deciding between buying an application and building on an API, our Dragon Medical alternatives guide covers that decision directly. For the broader clinical picture, see AI scribes and our medical solutions overview.

Working Through A Clinical Deployment

Talk through Medical Mode accuracy, BAA terms, EU data residency and self-hosted options with someone who has run these deployments before.

Talk to AI expert

Frequently asked questions

What is the most accurate medical speech recognition software?

Measure on Missed Entity Rate — the metric that counts dropped drugs, dosages and conditions rather than all words equally — rather than on word error rate. AssemblyAI’s Medical Mode posts a 3.2% MER, which is 87% fewer entity errors than the base model. Speechmatics is the strongest competitor on the medical entity row of our published benchmark. Methodology and audio sets are published at assemblyai.com/benchmarks, and you should re-run the comparison on your own clinical audio before deciding.

Why measure Missed Entity Rate instead of word error rate?

Word error rate treats every word as equally important, which is false in a clinical transcript: missing “the” is noise, missing “metoprolol” is a patient safety event. Missed Entity Rate isolates medical entities — drugs, dosages, conditions, procedures, anatomy — and measures how often they are lost. Two systems with nearly identical WER can differ substantially on MER, which is why general transcription benchmarks are a poor proxy for clinical performance.

How much does medical speech recognition cost?

Developer APIs bill per hour of audio: on AssemblyAI, $0.21/hr for pre-recorded transcription on Universal-3.5 Pro, $0.45/hr for streaming, and $0.62/hr for the Dictation API, with Medical Mode adding $0.15/hr — $0.36/hr combined with pre-recorded and $0.60/hr combined with streaming. Clinician-facing applications price differently, typically $49–99 per clinician per month, because they include the whole application. Compare within a category, not across the two.

Can medical speech recognition run on-premises or in my own cloud?

Yes, with different trade-offs. NVIDIA Riva is self-hosted by design and you operate the models yourself. AssemblyAI offers a self-hosted deployment — for example inside a customer’s own AWS VPC — plus EU data residency at the same price, with data staying in the EU. Speechmatics and Deepgram publish on-premises or self-hosted options as well. Confirm which deployment is available on the plan you will actually be on.

Which languages are supported for medical transcription?

Medical Mode supports English, Spanish, German and French across both pre-recorded and streaming. The underlying Universal-3.5 Pro model covers 18 languages with native code-switching on the streaming and async paths, and the Dictation API accepts 32 language codes. Clinical tuning and general language coverage are separate questions — a provider supporting 90 languages generally may still offer clinical tuning in only a handful, so verify both.

Does AssemblyAI sign a Business Associate Addendum?

Yes. AssemblyAI enables covered entities and their business associates subject to HIPAA to use the AssemblyAI services to process protected health information (PHI). AssemblyAI is considered a business associate under HIPAA, and we offer a standard Business Associate Addendum (BAA) that is required under HIPAA to ensure that AssemblyAI appropriately safeguards PHI. It can be reviewed and signed self-serve from the Data Controls page in the dashboard, and supporting controls include PHI redaction across audio and transcripts, SOC 2 Type 2, ISO 27001:2022 and PCI DSS v4.0.

Can I use the Dictation API for medical dictation?

Yes, for short-form clinical dictation — a chart note, an order, a message to a colleague — using stt_prompt to describe the setting, keyterms_prompt for drug names, and llm_instruction to shape the output as a chart note. Two limits matter: audio is capped at 120 seconds per call, which covers short dictations and sits at the edge of a typical 250–400 word SOAP note, and Medical Mode is not available on the Dictation API — there is no domain parameter, and the same applies to the Sync API. For longer notes or anywhere a missed medical entity is unacceptable, route to Universal-3.5 Pro pre-recorded or streaming with Medical Mode, where the 3.2% MER figure comes from.

Title goes here

Lorem ipsum dolor sit amet, consectetur adipiscing elit, sed do eiusmod tempor incididunt ut labore et dolore magna aliqua. Ut enim ad minim veniam, quis nostrud exercitation ullamco laboris nisi ut aliquip ex ea commodo consequat. Duis aute irure dolor in reprehenderit in voluptate velit esse cillum dolore eu fugiat nulla pariatur.

Button Text
Medical
Healthcare