Insights & Use Cases
September 22, 2026

Medical transcription services: how to choose one, and when to build instead

Medical transcription services deliver accurate, HIPAA-compliant documentation with fast turnaround and expert support for healthcare providers and clinics.

Kelsey Foster
, 
Growth
Reviewed by
No items found.
Table of contents

If you are searching for medical transcription services, you are usually deciding between three things that get sold under the same name: a service (people or a vendor who returns finished documents), software (an application your clinicians use directly), or an API (transcription you build into your own product). The right answer depends on whether you are buying an outcome, a tool, or a component.

This guide separates the three, prices each one, and shows what accuracy actually means in a clinical context. If you are building the product rather than buying it, the short version is at the bottom of every section: AssemblyAI’s Medical Mode runs on Universal-3.5 Pro at $0.36/hr all-in for pre-recorded audio and $0.60/hr for streaming, posts a 3.2% Missed Entity Rate, and comes with a Business Associate Addendum you can sign without a sales call.

What a medical transcription service actually delivers

A medical transcription service takes clinical audio — a dictated chart note, a recorded encounter, an operative report — and returns structured text that a clinician signs and an EHR stores. The differences between vendors come down to four things:

  • Who does the work. A transcriptionist, a model, or a model with a human reviewing the output.
  • Turnaround. Minutes, hours, or next business day.
  • What you get back. A raw transcript, a formatted note, or a note mapped into EHR fields.
  • How PHI is handled. Where audio lives, who can see it, and whether the vendor will sign a Business Associate Addendum.

Everything else is a variation on those four. The service models below are the standard shapes the market has settled into.

The three service models

Human medical transcription services

A clinician dictates; a trained medical transcriptionist listens and types. This is the original model and it still exists because it handles the cases automation is worst at: heavy accents, poor recording conditions, specialty vocabulary that shifts by department, and documents where a formatting convention matters as much as the words.

What you are buying is a finished document and someone accountable for it. What you are paying for is human time, which is why pricing is typically quoted per line, per minute of audio, or per report, and why turnaround is measured in hours rather than seconds. Cost scales linearly with volume — there is no point at which a hundred thousand notes get cheaper per note.

Choose it when: volume is low, documents are long and unusual, or the review burden of checking machine output would exceed the cost of having a person do it once.

AI medical transcription services

A model transcribes the audio and returns text with no human in the loop. Turnaround collapses from hours to seconds, and cost stops tracking headcount. Modern clinical speech models are not general-purpose transcription with a dictionary bolted on — they are tuned specifically to hold onto the tokens that carry clinical meaning: drug names, dosages, anatomy, procedure terms, negations.

The honest limitation is that an AI service returns output, not accountability. If a drug name is dropped, the model does not know it dropped it. That is why the metric that matters in clinical transcription is not word error rate but Missed Entity Rate — how often a clinically significant term goes missing, which is a very different question from how often a word is wrong.

Choose it when: volume is high, turnaround matters, and you have a review step somewhere downstream — usually the clinician signing the note.

Hybrid services (AI draft, human review)

The model produces a first draft in seconds; a human editor corrects and formats it before it reaches the clinician. This is where most of the market has actually landed, because it inverts the economics of the human-only model: the transcriptionist stops typing and starts editing, which is several times faster per document.

Hybrid pricing usually sits between the two, and turnaround sits between the two. The interesting variable is the edit rate — what fraction of documents need meaningful correction. A model with a low Missed Entity Rate lowers the edit rate, which is the single biggest lever on a hybrid service’s unit cost. This is why service vendors care about the underlying model’s medical entity accuracy even though their customers never see it.

Choose it when: you need machine speed and human accountability at the same time, which describes most health systems.

Comparing the three

Human AI Hybrid
Turnaround Hours to next day Seconds Minutes to hours
Cost behavior Scales with headcount Scales with audio volume Between the two; falls as model accuracy rises
Accountability Named reviewer Downstream signer only Named reviewer
Handles unusual audio Best Depends on the model Good — the editor catches what the model misses
Marginal cost at scale Flat Lowest Falls with edit rate
Typical buyer Small practice, specialty clinic Product team, high-volume operation Health system, transcription vendor

Services vs. software vs. API: the decision that actually matters

Most buyers arrive looking for a service and leave having bought something else, because the three categories solve different problems.

Service Software API
What you buy Finished documents An application clinicians use A component you build with
Who operates it The vendor Your clinicians Your engineers
Integration work None to minimal Configuration and rollout You own it end to end
Unit economics Per line / per report Per seat per month Per hour of audio
Control over output format Vendor templates Vendor templates Complete
Best fit You need documents, not a project You need a tool in clinicians’ hands You are shipping a product

Three questions resolve it almost every time:

  1. Are you the one billing for documentation, or consuming it? If you sell a clinical product, you need an API. If you run a clinic, you need a service or software.
  2. Does the output need to live inside an experience you control? Ambient scribes, clinical assistants, and specialty EHR modules all fail if transcription is a separate portal. That is an API decision.
  3. Is your volume predictable? Per-seat software is priced for steady use. Per-hour API pricing is priced for variable use and gets cheaper as you grow.

If you landed on “API,” the rest of this guide is about what to evaluate.

Start Free On Your Own Clinical Audio

Get 185 hours of pre-recorded and 333 hours of streaming transcription on the free tier, no card required.

Sign up free

How to evaluate clinical transcription accuracy

Word error rate is the number every vendor publishes and the wrong number to buy on. WER treats every token identically: dropping “the” and dropping “metoprolol” cost the same. In a clinical note they do not.

Missed Entity Rate (MER) measures how often a clinically significant entity — a medication, a dosage, an anatomical term, a procedure — fails to appear in the transcript. It is the metric that predicts whether a note is safe to sign. We wrote up the distinction in detail in MER vs. WER, and the underlying numbers for every provider we test are published on the benchmarks page.

Two other things are worth testing before you commit:

  • Speaker separation. A two-person encounter transcribed as one block of text is not usable for a note. Diarization quality is best measured by cpWER, which penalizes attributing the right words to the wrong speaker. Averaged across DiPCo, CALLHOME, NOTSOFAR and AMI, Universal-3.5 Pro posts cpWER 30.17, against Azure at 30.35, ElevenLabs Scribe V2 at 35.26, Speechmatics at 36.60, Gladia at 36.88 and Deepgram at 37.93. Azure is close enough that you should run both on your own encounter audio rather than deciding on the published average.
  • Contextual grounding. A model that knows what encounter it is transcribing performs measurably better than one that does not. See the section below.

Where AssemblyAI fits: Medical Mode on Universal-3.5 Pro

AssemblyAI is a Voice AI infrastructure platform, which places it firmly in the API column. Medical Mode is a single parameter on the flagship model rather than a separate product or endpoint.

Activation is one field. On pre-recorded audio:

import assemblyai as aai

aai.settings.api_key = "<YOUR_API_KEY>"

config = aai.TranscriptionConfig(
    speech_models=["universal-3-5-pro"],   # note: plural on async
    domain="medical-v1",
)

transcript = aai.Transcriber().transcribe("encounter.wav", config)
print(transcript.text)

On streaming the equivalent field is speech_model, singular, with the same domain: “medical-v1”.

What Medical Mode delivers

  • 3.2% Missed Entity Rate in absolute terms.
  • 87% fewer entity errors than the base model.
  • ~20% fewer missed medical entities than the base model.
  • English, Spanish, German, and French, on both pre-recorded and streaming paths.

What it costs

Path Base Medical Mode Combined
Pre-recorded (Universal-3.5 Pro) $0.21/hr +$0.15/hr $0.36/hr
Streaming (Universal-3.5 Pro Realtime) $0.45/hr +$0.15/hr $0.60/hr

Those are the full rates for the transcription itself. Optional add-ons are itemized on the pricing page — async diarization is +$0.02/hr, streaming diarization +$0.12/hr, and async keyterms prompting +$0.05/hr on Universal-3.5 Pro. Streaming keyterms are included.

Contextual prompting: the second accuracy lever

Medical Mode tunes the model for clinical vocabulary in general. Contextual prompting tells it about this specific encounter — the patient’s medication list, the specialty, the reason for the visit. They are complementary, and together they outperform either alone.

On a public benchmark of 20,000 real voice-agent calls, detailed context cut medical-term entity errors by 43%, and scenario-level context by 24% (prompting and keyterms documentation). In internal testing, feeding a patient’s prior-visit note cut missed medical terms by a further 31%.

The practical implication for a service model: if you are running a hybrid operation, contextual prompting is the cheapest available reduction in your edit rate.

Dictation methods

“Medical transcription” covers two genuinely different capture modes, and the service you need depends on which one you are in.

Phone-based dictation

The clinician dials a number, dictates, and hangs up. Audio lands in a queue for transcription. This is still common in hospitals because it requires nothing from the clinician beyond a phone, works from anywhere, and does not depend on a device being charged. The tradeoff is telephony-grade audio — narrowband, compressed, often noisy — which is the hardest input for any transcription model and the case where human or hybrid services still earn their cost.

Handheld dictation devices

A dedicated recorder or a smartphone app captures audio locally and uploads it in a batch. Audio quality is much better than phone capture, and the clinician can dictate without connectivity. The cost is a device to manage and a sync step that can fail silently. For product teams, handheld capture usually resolves to a pre-recorded transcription call — the pre-recorded STT path with Medical Mode enabled.

The Dictation API: short-form dictation that returns what the clinician meant

The third method is dictation that goes straight into the interface the clinician is already typing in — a chart note field, an order, a message to a colleague. This is the case the Dictation API was built for, and it behaves differently from the two above.

Transcription models are verbatim by design. That is correct for a legal record and wrong for text that is about to be sent. Speakers restate, self-correct, and hedge; what was said is not what anyone wants to send. The Dictation API runs speech-to-text and an LLM cleanup pass in a single low-latency call and returns both results:

  • $0.62/hr flat — everything included, one line on the bill, no token math.
  • 0.36 s to finished text, with typical short clips coming back in under a second.
  • stt_prompt — describe the situation in plain English, e.g. “A doctor dictating a patient visit note.”
  • keyterms_prompt — up to 100 exact terms to bias toward, which is where a medication formulary or a clinician’s patient panel goes.
  • llm_instruction — the rewrite task, in plain English. “Rewrite as a concise clinical chart note” is a valid instruction.
  • Both outputs returned. text is the verbatim transcript and is never altered by the LLM; llm_response is the cleaned version. Nothing is lost, so you can store the verbatim record and surface the clean one.
import assemblyai as aai

aai.settings.api_key = "<YOUR_API_KEY>"

config = aai.DictationConfig(
    stt_prompt="A doctor dictating a patient visit note.",
    keyterms_prompt=["amoxicillin", "lisinopril", "metoprolol"],
    llm_instruction=(
        "Remove filler words and rewrite as a concise clinical chart note."
    ),
)

result = aai.DictationTranscriber().transcribe_live("clip.wav", config)

transcript = result.text
rewrite = result.final_text  # the cleaned-up text, falling back to the transcript

Two limits to plan around

The cap is 120 seconds of audio per call. A routine SOAP note runs roughly 250–400 words, which is about two to three minutes at conversational dictation speed — so 120 seconds covers short dictations and sits right at the edge of a typical note. Complex hospitalist or psychiatric notes run 800+ words and will not fit. For those, route to pre-recorded STT or the streaming path instead.

Medical Mode is not available on the Dictation API. domain: “medical-v1” runs on Universal-3.5 Pro async and Universal-3.5 Pro Realtime only. There is no domain parameter on the Dictation endpoint. If the 3.2% MER figure is what you are underwriting your accuracy story on, you need the async or streaming path — not the Dictation API. The Dictation API still runs on Universal-3.5 Pro and still takes keyterms_prompt, which covers a lot of clinical vocabulary in practice, but it is not the same thing and we are not going to pretend it is.

One more honest note for this buyer: clinicians are trained to speak punctuation — “Medications colon aspirin comma naproxen period.” AssemblyAI’s models apply punctuation automatically and handle spoken command words inconsistently, and there is no verbatim/spoken-form mode today. Deepgram ships a dictation=true flag that converts spoken commands into punctuation and markets it in their medical content; if that behavior is a hard requirement, you should know it exists. The documented workaround on our side is an LLM pass over each finalized turn to interpret and strip command words, which is the pattern in the medical scribe best practices guide. It works, and it is a workaround rather than a native mode.

Try It Before You Build

Run clinical audio through Medical Mode and the Dictation API in the playground without writing any code.

Try playground

PHI, security, and what to ask a vendor

AssemblyAI is a business associate under HIPAA, not a covered entity. The legal-approved description:

AssemblyAI enables covered entities and their business associates subject to HIPAA to use the AssemblyAI services to process protected health information (PHI). AssemblyAI is considered a business associate under HIPAA, and we offer a standard Business Associate Addendum (BAA) that is required under HIPAA to ensure that AssemblyAI appropriately safeguards PHI.

The practical detail buyers care about: the BAA can be reviewed and signed self-serve from the Data Controls page in the dashboard, without a sales call. Details are at can you sign a BAA and the Business Associate Addendum page.

Supporting controls, none of which are a substitute for the BAA itself: PHI redaction across audio and transcripts, SOC 2 Type 2, ISO 27001:2022, and PCI DSS v4.0. Deployment options include hosted Voice AI Cloud, self-hosted in your own cloud account, and EU data residency where data stays in the EU at the same price. The full list is on the security page.

When you evaluate any transcription service — human, AI, or hybrid — ask the same four questions: will you sign a BAA, where does audio reside, who on your side can access it, and how long is it retained.

Putting it together

If you run a practice and need documents, buy a service, and pick human or hybrid based on how unusual your audio is. If you need a tool clinicians open every day, buy software. If you are building a clinical product — an ambient scribe, an assistant, a specialty EHR module — buy an API, and evaluate it on Missed Entity Rate and diarization quality rather than word error rate.

For the API path, the shape of the decision is straightforward: pre-recorded encounters go to Universal-3.5 Pro with Medical Mode at $0.36/hr; live capture goes to Universal-3.5 Pro Realtime with Medical Mode at $0.60/hr; short-form front-end dictation goes to the Dictation API at $0.62/hr, with the 120-second cap and the absence of Medical Mode understood going in.

Map Your Workflow To The Right Surface

Talk through Medical Mode, BAA terms, EU data residency and self-hosted options with someone who has run these deployments before.

Talk to AI expert

Frequently asked questions

How much do medical transcription services cost?

Human and hybrid services are priced per line, per minute of audio, or per report, and cost scales with the amount of human review involved. API-based transcription is priced per hour of audio: AssemblyAI charges $0.21/hr for pre-recorded Universal-3.5 Pro and $0.45/hr for streaming, with Medical Mode adding $0.15/hr to either — $0.36/hr and $0.60/hr respectively. The Dictation API is $0.62/hr flat with every feature included.

How accurate is AI medical transcription?

Measured on the metric that matters clinically, AssemblyAI’s Medical Mode posts a 3.2% Missed Entity Rate — that is, 3.2% of clinically significant entities are missed. That represents 87% fewer entity errors than the base model. Word error rate is the more commonly published number but it weights “the” and “metoprolol” identically, so it is the wrong metric for clinical audio. Run the comparison on your own audio before deciding; published benchmarks narrow a field, they do not settle it.

Is AssemblyAI HIPAA-compliant?

AssemblyAI is a business associate under HIPAA and offers a standard Business Associate Addendum (BAA), which is what HIPAA requires for a vendor processing PHI on your behalf. The BAA can be reviewed and signed self-serve, without a sales call. Supporting certifications include SOC 2 Type 2, ISO 27001:2022, and PCI DSS v4.0, and PHI redaction is available across both audio and transcripts.

What is the difference between medical transcription and an ambient AI scribe?

Medical transcription converts dictated or recorded audio into text. An ambient AI scribe listens to a live clinician-patient conversation and produces a structured note from it, which requires speaker separation, far-field audio handling, and a summarization step on top of transcription. Transcription is a component of an ambient scribe, not a substitute for one.

Can I transcribe medical audio in languages other than English?

Medical Mode supports English, Spanish, German, and French on both pre-recorded and streaming paths. The underlying Universal-3.5 Pro model covers 18 languages with native code-switching, and Universal-2 covers 99+ languages for legacy integrations — but the medical tuning itself is limited to those four.

How fast is AI medical transcription?

Pre-recorded transcription returns in a fraction of the audio duration. Streaming returns transcripts continuously as the audio arrives, which is fast enough for live display during an encounter; benchmark the end-to-end figure on your own audio rather than relying on a published headline, because the number vendors quote often describes a different span than the one your clinicians experience. The Dictation API returns finished, cleaned-up text in about 0.36 s for a short clip. Human and hybrid services, by comparison, are measured in hours.

Is AI replacing medical transcriptionists?

No — the role is shifting rather than disappearing. As model accuracy improves, the work moves from typing a document to reviewing, correcting, and formatting one, which is what hybrid services already run on today: a model drafts in seconds, a human editor validates. What is genuinely changing is the mix of the job. Volume transcription of clean audio is increasingly automated, while the human work concentrates in quality assurance, exception handling for difficult audio and unusual specialties, and the final accountability step before a note is signed. Editing throughput per person is much higher than typing throughput, so the same team covers far more documents — which is a change in what the job is, not an elimination of it.

Title goes here

Lorem ipsum dolor sit amet, consectetur adipiscing elit, sed do eiusmod tempor incididunt ut labore et dolore magna aliqua. Ut enim ad minim veniam, quis nostrud exercitation ullamco laboris nisi ut aliquip ex ea commodo consequat. Duis aute irure dolor in reprehenderit in voluptate velit esse cillum dolore eu fugiat nulla pariatur.

Button Text
Medical