New Universal-3.6 Pro Realtime is now available Learn more
Insights & Use Cases

Conversational AI in healthcare: maturity model and 7 use cases

Learn how conversational AI in healthcare is driving high-impact use cases for everything from ambient documentation to mental health support.

Abstract green diamond illustration

Written by

Jesse Sumrak

Published on

30 September 2026

Picture a health system's IVR. It routes every inbound call through a menu tree that ends, for a large share of callers, in a queue. The system works exactly as designed and nobody likes it — not the caller who presses 4 twice and then waits, and not the scheduler who spends the afternoon on calls a machine could have closed in ninety seconds.

Conversational AI is what replaces that, and the reason it's finally viable in healthcare isn't better dialogue design. It's that the speech layer got good enough at clinical vocabulary to be trusted with medication names. Medical Mode on Universal-3.5 Pro posts a 3.2% Missed Entity Rate — the lowest across benchmarked providers, against roughly 8.7% for Deepgram and roughly 24.4% for AWS Transcribe Medical.

That threshold is what separates a demo from a deployment. Below the threshold, every conversational system in healthcare is a liability that needs a human reading every transcript. Above it, you can automate.

This post covers what conversational AI in healthcare actually is, how it differs from the rule-based systems it's replacing, a four-stage maturity model for figuring out where you are, seven use cases with the specific problem each one solves, and what it costs.

What conversational AI in healthcare means

Conversational AI is any system that holds a natural-language exchange with a patient, member, or clinician and takes action from it. In healthcare it shows up in four places: inbound and outbound phone calls, patient portal chat, clinical documentation from ambient conversation, and post-visit or post-discharge follow-up.

What unifies them is a dependency most buyers underweight. Every one of these systems is built on top of a transcript. If the transcript is wrong, the intent classification is wrong, the routing is wrong, the summary is wrong, and the chart entry is wrong. Conversational AI in healthcare is a speech problem wearing a dialogue problem's clothes.

Advanced voice AI versus rule-based systems

The systems being replaced aren't stupid. They're rigid, which is worse.

Dimension Rule-based IVR / scripted bot Voice AI on a modern speech model
Input handling Keypad or a fixed grammar of accepted phrases Open speech, however the caller phrases it
Clinical vocabulary Only terms someone enumerated in advance Domain-trained on drugs, conditions, procedures
Multiple speakers Breaks Diarization, up to 10 speakers with revision
Language switching Separate menu per language, chosen up front Native code-switching across 32 languages
Changing the flow A ticket, a vendor, a release window A prompt and a config change
What you learn from a call Which buttons were pressed Full transcript, entities, speaker attribution

That last row is the one I'd put in front of a CFO. A rule-based system produces almost no data about why people are calling. A voice AI system produces a searchable transcript of every conversation your organization has with a patient. The automation is the near-term win; the corpus is the durable one.

Where the business case actually comes from

I'd be skeptical of anyone offering you a single ROI multiple for this. The value shows up in four different places and they don't share a denominator.

Deflected calls. Every reschedule, refill status check, and eligibility question closed without a human is measurable directly against your call center's cost per contact.

Recovered clinician time. Ambient documentation moves note-writing out of the evening. This is the benefit clinicians feel most, and it's the one that shows up in retention rather than in a cost line.

Coverage you didn't have. The 7:40 p.m. caller who currently leaves a voicemail. That's not a cost saving; it's revenue and access that didn't exist.

Fewer downstream corrections. This is where entity accuracy converts to money. If a missed medication name means a note goes back for correction, then cutting missed entities cuts rework. Medical Mode delivers roughly 20% fewer missed medical entities and 87% fewer entity errors versus the base model without it. Those aren't abstract quality metrics — they're the review queue getting shorter.

Model the four separately. A blended number will be wrong and someone will catch it.

Run The Accuracy Test Before You Model The ROI

Sign up free and transcribe your own call recordings and encounter audio. The entity accuracy you measure on your own data is the input every other number depends on.

Sign up free

A four-stage maturity model

Most organizations are further back than they think, and the failure mode is skipping stages.

Stage 1: foundational — capture and transcribe

You're recording calls and encounters and getting accurate transcripts out. No automation yet. This stage sounds trivial and it's the one people rush. Get transcript quality right here, measured on your own audio with your own vocabulary, or every later stage inherits the error.

What to have working: async transcription at $0.36/hr with Medical Mode, speaker labels, PHI redaction, a BAA in place, and a retention policy.

Stage 2: operational — search, route, and summarize

Transcripts become useful. You can search across calls, classify intent after the fact, auto-summarize encounters, and pull structured fields out of conversations. Still no live automation, but the corpus is now doing work.

Stage 3: interactive — live agents on bounded tasks

Real-time systems handle specific, verifiable intents end to end. Scheduling. Refill status. Eligibility. Universal-3.6 Pro Realtime at $0.60/hr with Medical Mode, final transcripts a median 307 ms after the speaker stops on Pipecat's open benchmark, and a fast escalation path. See voice agents in healthcare for the build detail.

Stage 4: transformative — conversation as clinical input

The transcript stops being a record and becomes a signal. Ambient documentation that drafts the note and pre-populates the order set. Population-level analysis across every patient conversation. Live prompts that catch a missed screening question. This stage requires the accuracy floor from stage 1 to be genuinely solid, because the system is now acting rather than reporting.

My honest read: the majority of healthcare organizations are at stage 1 or 2 and buying stage 3 pilots. That's how you get a voice agent that impresses in a demo and gets switched off in month four.

Seven use cases and the problem each one solves

1. Appointment scheduling and rescheduling

The problem: schedulers spend their day on calls with a single, bounded outcome, and callers who reach a queue give up. The solution: a live agent that reads availability, confirms, and writes back. This is the highest-confidence starting point in healthcare.

2. Ambient clinical documentation

The problem: clinicians write notes after hours. The solution: transcribe the encounter with speaker labels and Medical Mode, then structure the note with an LLM. Accuracy on drug names is the whole game — see the AI medical scribe build guide.

3. Prescription refill and pharmacy routing

The problem: high-volume, low-complexity calls that still require exact medication capture. The solution: an agent that confirms drug, dose, and pharmacy — with Medical Mode on, because "Lamictal" and "Lamisil" are one phoneme apart and treat entirely different things.

4. Insurance eligibility and benefits

The problem: rules-based questions that consume enormous staff time. The solution: full automation with escalation on anything unusual. Nothing clinical happens here, which makes it low-risk and high-volume — a good second deployment.

5. Post-discharge follow-up

The problem: follow-up calls are the ones that get dropped when staffing is tight, and they're the ones that catch readmissions early. The solution: outbound structured check-ins with immediate escalation on concerning answers. The agent asks; a human assesses.

6. Multilingual patient access

The problem: patients whose language isn't the one your phone tree opens in, and calls where a family member interprets mid-sentence. The solution: two capabilities, held separately. Medical Mode covers English, Spanish, German, and French. The base model code-switches natively across 18 languages for pre-recorded audio, and Universal-3.6 Pro Realtime covers 32 languages for streaming with automatic language detection and no configuration, which is what handles the mid-sentence flip.

7. Call center quality and intent analytics

The problem: you don't know why people call, only which menu options they chose. The solution: transcribe every call async at $0.21/hr, classify intent, and find out. This is the cheapest project on the list and frequently the most immediately useful — it tells you which of the other six to build.

How to actually implement it

The pattern that works, in order:

Start with async on existing recordings. No live system, no patient exposure, no risk. Transcribe a few thousand real calls or encounters and measure entity accuracy on your own vocabulary. You'll learn more in two weeks than in six months of vendor evaluation.

import assemblyai as aai

aai.settings.api_key = "YOUR_API_KEY"

config = aai.TranscriptionConfig(
    speech_models=["universal-3-5-pro"],
    domain="medical-v1",
    speaker_labels=True,
    redact_pii=True,
    redact_pii_audio=True, redact_pii_policies=[aai.PIIRedactionPolicy.person_name, aai.PIIRedactionPolicy.date_of_birth, aai.PIIRedactionPolicy.phone_number],
)

transcript = aai.Transcriber().transcribe("call-archive/0042.wav", config=config)
print(transcript.text)

Add context. Contextual prompting is the most underused accuracy lever available. Feeding a patient's prior-visit note as context cut missed medical terms by 31% in an internal healthcare test. Your formulary, your provider names, your specialty vocabulary — the model performs better when you tell it what to expect.

Then go live on one intent. One. Scheduling. Instrument escalation rate and entity accuracy per call. Expand only when both hold.

Keep the async pipeline running underneath. Re-transcribe live calls after the fact with the async model. It sees the whole recording, so it catches what the live pipeline missed, and that diff is your QA signal.

See Medical Mode Against Your Own Vocabulary

Drop a call recording into the playground, toggle Medical Mode, and compare the medication names side by side. Two minutes, no integration.

Try playground

What the platform provides

The pieces that matter for healthcare conversational AI, concretely:

Async transcription. Universal-3.5 Pro, speech_models: ["universal-3-5-pro"], $0.21/hr, or $0.36/hr with Medical Mode.

Streaming. Universal-3.6 Pro Realtime over wss://streaming.assemblyai.com/v3/ws, $0.45/hr base or $0.60/hr with Medical Mode, with agent_context, voice_focus, and latency modes from min_latency to max_accuracy.

Voice Agent API. One WebSocket replacing speech-to-text, LLM, and text-to-speech at a flat $4.50/hr.

Diarization. The most accurate AssemblyAI has shipped, optimized for cpWER rather than DER, holding up on short turns and overlapping speech — which is what a doctor-patient conversation is.

PHI handling. AssemblyAI signs a Business Associate Addendum (BAA) for customers processing PHI and acts as a business associate under HIPAA. SOC 2 Type 2. PHI redaction across both audio and transcripts. EU data residency at api.eu.assemblyai.com and a self-hosted option. See the BAA FAQ and the BAA page.

Teams working in this space on AssemblyAI include Sully AI, Heidi Health, Magentus Healthcare, Deepscribe, Knowtex, Chapter, NMDP, and Commure. Full rates are on pricing; parameter reference in the docs.

What changes when every conversation is searchable

The automation story is the one that gets budget, but it's not the interesting one. The interesting one is what happens to an organization that has, for the first time, an accurate, structured, searchable record of every conversation it has with patients.

You can answer questions that were previously unanswerable. Which medication is most frequently confused on our refill line, and does that correlate with a particular pharmacy? Are Spanish-speaking patients escalating to humans at a higher rate, and at which step? Which screening questions get skipped in which clinics? Did the new discharge instruction script actually change what patients say back to us?

None of those need a voice agent. They need transcripts good enough to trust and a decision to look. My prediction is that the organizations that get the most out of conversational AI over the next few years won't be the ones with the most automation — they'll be the ones who treated stage 1 as the strategic investment it is, built the corpus, and then discovered that the answers were sitting in their own call archive the whole time.

Plan Your Stage-By-Stage Rollout

Our team works with health systems and health tech companies on accuracy evaluation, BAA scope, data residency, and which use case to automate first.

Talk to AI expert

Frequently asked questions

What is AssemblyAI's Medical Mode?

Medical Mode is a domain setting for AssemblyAI's speech models, activated with a single parameter — domain: "medical-v1" — with no model switch required. It tunes recognition for clinical vocabulary: drug names, conditions, procedures, anatomy, dosages. It costs $0.15/hr on top of the base model, so $0.36/hr combined for pre-recorded audio and $0.60/hr for streaming, and it works on both Universal-3.5 Pro and Universal-3.6 Pro Realtime.

How accurate is AssemblyAI Medical Mode compared to other providers?

Universal-3.5 Pro with Medical Mode records a 3.2% Missed Entity Rate on medical entities — the lowest across benchmarked providers. Deepgram comes in around 8.7% MER and AWS Transcribe Medical around 24.4% MER. Against the base model without Medical Mode, it captures roughly 20% fewer missed medical entities and 87% fewer entity errors. Methodology is on the benchmarks page.

How does AssemblyAI handle HIPAA and PHI?

AssemblyAI signs a Business Associate Addendum (BAA) for customers processing PHI and operates as a business associate under HIPAA. Supporting controls include SOC 2 Type 2, PHI redaction across both audio and transcripts, EU data residency, and self-hosted deployment for teams that can't send audio outside their own environment. Start with the BAA FAQ.

Does AssemblyAI automatically redact patient PII from medical transcripts?

Yes, with a parameter. PII and PHI redaction runs across both the transcript and the audio file itself, so identifiers can be removed from the recording as well as the text. Turn it on with redact_pii and redact_pii_audio in your transcription config. The docs list the entity categories you can target.

Which languages does conversational AI for healthcare support?

Two separate things. Medical Mode covers English, Spanish, German, and French for both pre-recorded and streaming audio. The base model independently code-switches natively across 18 languages for pre-recorded audio and 32 languages on Universal-3.6 Pro Realtime for streaming, with no configuration, which handles calls where the caller changes language mid-sentence — the common case when a family member interprets.

How does conversational AI connect to an EHR?

AssemblyAI produces the transcript, speaker labels, entities, and word-level timings; your application writes to the chart through the EHR's own interface — typically FHIR, HL7, or a vendor API. There's no direct EHR connector, which is deliberate: the transcript layer stays vendor-neutral so you're not re-platforming your speech stack when you change EHRs. See the medical transcription use case for reference architecture.