New Universal-3.6 Pro Realtime is now available Learn more
Insights & Use Cases

How to build an AI medical scribe with AssemblyAI

Build a production-ready AI medical scribe with Python. Learn speaker identification, PII redaction, SOAP note generation, and automatic data deletion.

Abstract green cylinder illustration

Written by

Martin Schweiger

Published on

30 September 2026

A working AI medical scribe demo takes an afternoon. Audio in, transcript out, LLM writes a SOAP note, everyone in the room nods. The version a clinician will actually sign their name to takes considerably longer, and the difference isn't the model — it's the six or seven things around the model that nobody demos.

The one that bites hardest is entity accuracy. A transcript can post an impressive word error rate while dropping every drug name in the encounter, because drug names are a rounding error in total word count. The metric that catches this is Missed Entity Rate, and it's where Medical Mode earns its keep: Universal-3.5 Pro with domain: "medical-v1" posts a 3.2% MER, the lowest across the providers we've benchmarked.

This is the build guide: what to put in, in what order, with the code. Python throughout. I'll be direct about the parts where the right answer is a design decision rather than an API call.

What you're actually building

An AI medical scribe listens to a clinical encounter and produces a note the clinician reviews and signs. Five stages, and each one can independently ruin the output:

  1. Capture — audio from a room microphone, a phone, or a telehealth stream.
  2. Transcribe and diarize — words plus who said them.
  3. Contextualize — what does the system already know about this patient?
  4. Structure — conversation into a SOAP note, HPI, assessment and plan.
  5. Govern — redaction, retention, deletion, audit.

Most teams build 1, 2, and 4, ship it, and then spend six months retrofitting 3 and 5. Build all five from the start; stage 5 in particular is much harder to add after you have production PHI sitting in a database.

The feature checklist worth arguing about

Requirement Why it’s non-negotiable How you get it
Medical entity accuracy A missed drug name is a clinical error, not a typo domain: "medical-v1"
Speaker attribution “Patient reports” vs “clinician observed” are different records speaker_labels: true
Patient context Cuts missed medical terms 31% in internal testing Contextual prompting with the prior note
Grounded note generation A hallucinated dose is worse than a blank field Schema + span citation + review flags
PHI redaction You’ll want de-identified copies for analytics and QA Redaction across audio and transcripts
Retention control and deletion Every hospital security review asks Explicit delete calls in your own pipeline
Consent capture Recording rules vary by state; no vendor supplies this Your product, your audit trail

Step 1: transcribe with Medical Mode on

Start here, because everything downstream inherits these errors. Async is the right default: it's cheaper, it sees the whole file before committing to anything, and post-visit notes don't need live output.

import assemblyai as aai

aai.settings.api_key = "YOUR_API_KEY"

config = aai.TranscriptionConfig(
    speech_models=["universal-3-5-pro"],
    domain="medical-v1",
    speaker_labels=True,
    speakers_expected=2,
    punctuate=True,
    format_text=True,
)

transcript = aai.Transcriber(config=config).transcribe(
    "./encounter.wav"
)

if transcript.status == "error":
    raise RuntimeError(transcript.error)

Two easy mistakes. speech_models is plural and takes an array for async requests — streaming uses the singular speech_model. And check status before you touch anything: a silently failed job that flows into note generation is how a chart gets a note about nothing.

What the domain flag buys you against the same model with it off: roughly 20% fewer missed medical entities and 87% fewer entity errors. It costs $0.15/hr on top of the $0.21/hr base, so $0.36/hr, plus $0.02/hr for diarization. It covers English, Spanish, German, and French. Methodology for the benchmark numbers is on the benchmarks page.

Step 2: get the speakers right

Diarization deserves more attention than it usually gets, because a scribe consumes speaker-attributed text and a wrong attribution produces a confidently wrong note.

Clinical encounters are hostile to standard diarization. The clinician talks in paragraphs; the patient answers in two or three words, frequently over the top of the question. Diarizers tuned for clean turn-taking drop exactly those short turns, and the patient's short turns are often the clinically important content.

Universal-3.5 Pro's diarization is the most accurate we've shipped, and it's optimized for cpWER rather than DER. That's not a technicality: DER measures how much audio time landed on the right speaker, which barely penalizes losing a two-word answer, while cpWER measures the words in each speaker's transcript — exactly what your note generator reads.

for utterance in transcript.utterances:
    print(f"Speaker {utterance.speaker}: {utterance.text}")

You'll get A, B, and so on. Mapping those to clinician and patient is your job. The reliable heuristics: whoever speaks first and longest is usually the clinician, and the speaker asking most of the questions almost always is. Don't guess silently — if confidence is low, surface it in the review UI rather than mislabeling the record.

See The Diarized Clinical Transcript First

Upload a real encounter, turn on Medical Mode and speaker labels, and read what your note generator would receive.

Try playground

Step 3: give the model context (the step most teams skip)

This is the single biggest accuracy win in the entire build, and it costs a database query.

In an internal healthcare test, feeding the model a patient's prior-visit note as context cut missed medical terms by 31%. No fine-tune, no new model. Just handing over the document that already lists this patient's medications, conditions, and specialists before transcribing their next visit.

It makes sense once you think about it. The hardest terms in any encounter are proper nouns specific to this patient — the biologic they started last month, the surgeon they're seeing, the local abbreviation this clinic uses. A general clinical model can't know those. Your database does.

If you're building something conversational rather than a transcription feed — an intake bot, a triage line — the streaming equivalent is agent_context, which cut word error rate 10.2% across 20,000 voice agent files, with detailed context cutting medical-term entity errors 43% on the same set. Details in the Universal-3.6 Pro Realtime post.

The architectural implication is worth stating plainly: stop maintaining specialty vocabulary lists. Fetch the patient's last note instead. It's less work and it works better.

Step 4: generate the note without letting the model invent

Now you have a clean, diarized, contextualized transcript. Turning it into a SOAP note is an LLM job, and it's where the interesting risk lives.

Three rules I'd enforce in code, not in a prompt:

Require an explicit schema

Subjective, objective, assessment, plan, medications, follow-up. Named fields, typed, validated on the way out. Free-text output invites drift.

Require span citation for every clinical claim

Every medication, dose, and diagnosis in the output should carry the transcript span it came from. If the model can't cite it, it doesn't go in the note — it goes in a review queue.

Prefer gaps over guesses

If the dosage wasn't stated aloud, the correct output is an empty field flagged for the clinician. A plausible invented number is the single worst failure mode this system has, because it looks exactly like a correct one.

Pass the transcript through the LLM Gateway with those constraints — the docs cover the request shape. And build the review UI as a first-class surface, not an afterthought. The metric that decides whether clinicians keep using your product is seconds spent editing per note.

Step 5: redact PHI

PHI redaction runs across both transcripts and audio, which matters more than it sounds. Redacting only text leaves you with a recording that still contains the patient's name — so you either can't keep the audio for QA, or you keep it under full PHI controls forever. Redacting both means you can hold a de-identified transcript and a redacted recording for model evaluation, debugging, and analytics.

Decide your redaction policy before you store anything. Retrofitting de-identification onto a database that already holds production PHI is a project, not a patch.

Step 6: retention, deletion, and audit

Your pipeline should be able to answer three questions on demand: what audio and transcripts exist for this patient, who accessed them, and how do I delete them right now.

Build explicit deletion into the pipeline rather than relying on a default. When an encounter's note is signed and stored in the EHR, the working transcript usually has no reason to exist anymore. Delete it on a schedule you can describe to a hospital security reviewer in one sentence.

Consent is the other piece nobody hands you. Recording clinical encounters carries state-level requirements, and the answer needs to be a documented step in your product with an audit trail — not a line in a terms-of-service document.

Start Building With Free Credits

No card required. The exact model IDs and parameters in this guide work on your first request.

Sign up free

Live scribing, if you need it

Everything above assumes post-visit processing. If your product shows a transcript during the encounter, add a streaming path at wss://streaming.assemblyai.com/v3/ws:

from assemblyai.streaming.v3 import (
    StreamingClient,
    StreamingClientOptions,
    StreamingParameters,
    TurnEvent,
)

def on_turn(client, event: TurnEvent):
    if event.end_of_turn:
        print(event.transcript)

client = StreamingClient(StreamingClientOptions(api_key="YOUR_API_KEY"))
client.on(TurnEvent, on_turn)

client.connect(
    StreamingParameters(
        speech_model="universal-3-6-pro",
        domain="medical-v1",
        sample_rate=16000,
        speaker_labels=True,
        voice_focus="far-field",
        mode="max_accuracy",
    )
)

Three settings to get right. voice_focus takes near-field for headsets and phones or far-field for a room microphone — ambient scribing is almost always far-field, and the default costs you accuracy silently. mode takes min_latency, balanced, or max_accuracy; documentation should sit at the accurate end, since nobody's waiting on a 200ms round trip to read a note about a visit that just ended. And streaming diarization supports revision, correcting early speaker guesses as more audio arrives, up to 10 speakers.

Streaming runs $0.45/hr base, $0.60/hr with Medical Mode. If what you're building is genuinely conversational — an intake agent that talks back — the Voice Agent API replaces your STT, LLM, and TTS stack with one WebSocket at a flat $4.50/hr. Current rates on the pricing page.

What it costs to run

Component Rate Per 15-minute encounter
Async transcription + Medical Mode $0.36/hr About nine cents
Streaming + Medical Mode $0.60/hr About fifteen cents
Voice Agent API $4.50/hr flat Conversational use only

Transcription is not your cost problem. Your real spend is EHR integration, the review interface, and clinician onboarding. Anyone telling you the speech layer is the expensive part hasn't built one.

PHI, BAAs, and the security review

AssemblyAI signs a Business Associate Addendum (BAA) for customers processing PHI, which makes us a business associate under HIPAA. The terms and request path are documented at can you sign a BAA and legal/business-associate-agreement — hand those to your compliance reviewer directly rather than paraphrasing them.

Supporting facts you'll be asked for: PHI redaction across audio and transcripts, SOC 2 Type 2, and self-hosted deployment or EU data residency via api.eu.assemblyai.com for teams with isolation or residency requirements. Teams shipping clinical documentation on this stack include Sully AI, Heidi Health, Deepscribe, Knowtex, Magentus Healthcare, and Commure.

What to build next

The version of this system worth building in 2027 isn't a better transcriber. It's a scribe that stops starting cold.

Right now, almost every AI medical scribe on the market hears one encounter with no memory of the last one. That's why the 31% figure from contextual prompting is the most important number in this post — it says the accuracy ceiling has moved out of the model and into your data layer. A scribe that reads the prior note before the visit is measurably better than one that doesn't, and a scribe that reads the last five, knows the clinic's shorthand, and learns which terms this particular physician actually uses will be better still.

That's a retrieval and data-modeling problem, which is good news: it's the kind of problem your team can solve without waiting on a model release. Build the context path early and the accuracy improvements keep arriving without you changing a single API call.

Related reading: AI medical transcription, medical voice recognition, the best ambient AI scribes if you're weighing build versus buy, the healthcare solutions overview, and speech-to-text.

Pressure-Test Your Scribe Architecture

Bring your context-retrieval plan, retention policy, and BAA requirements. Our team will tell you which of the six stages will bite you first.

Talk to AI expert

Frequently asked questions

What's the best API for building an AI medical scribe?

Judge on medical entity accuracy, diarization quality, and whether medical support is a parameter or a separate model. Universal-3.5 Pro with Medical Mode posts a 3.2% Missed Entity Rate — the lowest across the providers we've benchmarked — with cpWER-optimized diarization, and the same domain parameter works for async and for streaming on Universal-3.6 Pro Realtime. See the benchmarks page, then test on your own encounters.

How do I stop the note generator from hallucinating?

Constrain it in code rather than in the prompt. Require a typed output schema, require a transcript span citation for every medication, dose, and diagnosis, and route anything uncited into a review queue instead of the note. An empty field flagged for the clinician is safe; a plausible invented dose is the worst failure this system can produce because it's indistinguishable from a correct one.

How does AssemblyAI handle HIPAA and PHI?

AssemblyAI signs a Business Associate Addendum (BAA) for customers processing PHI, which makes us a business associate under HIPAA. Alongside the BAA you get PHI redaction across audio and transcripts, SOC 2 Type 2, and self-hosted or EU-residency deployment options. Start at can you sign a BAA and send the link to your compliance team.

Does AssemblyAI automatically redact patient PII from medical transcripts?

Yes, and across both the transcript and the audio — which is the part that matters. Text-only redaction leaves you holding a recording that still contains the patient's name, so you either delete the audio or keep it under full PHI controls. Redacting both lets you retain de-identified copies for QA and model evaluation. Set the policy before you store anything.

Should I build an AI medical scribe or buy one?

Build if documentation is part of a product you sell, if your workflow isn't a standard SOAP note, or if per-clinician subscription pricing breaks your economics — at $0.36/hr the transcription layer is close to free at scale. Buy if you're a provider organization documenting your own visits; finished products have already solved consent flows, EHR write-back, and clinician onboarding. We compare the options in the ambient AI scribes roundup.

How does an AI medical scribe handle multilingual encounters?

The base model code-switches natively across 18 languages with no configuration, so an encounter that shifts between English and Spanish mid-sentence transcribes correctly without language detection or request routing. Medical Mode's clinical entity accuracy covers English, Spanish, German, and French. Keep those two facts separate when you're scoping — they answer different questions about your patient population.