New Universal-3.5 Pro is here. Learn more: Async Realtime
Voice agents · Healthcare

Voice agents for healthcare and clinical workflows

In healthcare, transcription accuracy isn't a product feature — it's a patient safety requirement. A misheard drug name or dosage is a clinical risk, not a UX annoyance. Here's how to build ambient scribes, medical intake, and clinical voice agents on infrastructure built for it.

Medical error rate · Lower is better

Clinical terminology accuracy

AssemblyAI Medical ModeDeepgramSpeechmatics Enhanced MedicalDeepgram Nova-3 MedicalAWS Transcribe Medical

Medical entity error rate across providers, lower is better. Source: AssemblyAI benchmarks.

The numbers that matter

Fewer entity errors

87%

Medical Mode reduces medical entity errors ~87% versus the base model — the headline accuracy gain on clinical terminology.

Source: AssemblyAI benchmarks

Medical error rate

3.2%

Medical Mode Missed Entity Rate across specialties — oncology, cardiology, primary care — without retraining.

Source: AssemblyAI benchmarks

Compliance

BAA

Business Associate Agreement available at no additional cost, alongside SOC 2 Type 2, ISO 27001:2022, and PCI DSS v4.0.

Source: BAA FAQ

Self-hosted

48

Concurrent sessions per container for teams that need complete data isolation (Docker + NVIDIA GPU).

Source: Self-hosted deployment

Customer story · Report 3

“We've integrated the newest models from AssemblyAI for pre-recorded audio ASR in our ambient product, and it's been excellent. We're now exploring Universal-3.5 Pro for async and realtime speech-to-text capabilities for new use cases. What's been just as important is the reliability of the platform itself—both technically and in terms of partnership.”

— Gautam Pradeep, Tech Lead, Commure

A different standard

Why healthcare voice agents demand a different standard

Healthcare audio is categorically harder: far-field ambient capture across the room, overlapping speakers (provider, patient, staff, family), dense terminology (drug names, procedures, dosages), and the highest-consequence failure mode of any voice AI application. Generic voice agent APIs weren't built for this — and "good enough" transcription creates clinical documentation that's actively dangerous.

Purpose-built for clinical terminology

Medical Mode optimizes transcription for medication names, procedures, conditions, and dosages — 3.2% MER, ahead of every general and medical model tested. $0.15/hr add-on.

Medical Mode

Knows who is speaking

Speaker role identification (NURSE, PATIENT, PROVIDER) plus SpeakerRevision for provider/patient/family disambiguation in multi-party encounters.

Universal-3.5 Pro Realtime

PHI handled by default

Automatic PHI redaction (text and audio), keyterms prompting for patient-specific drug lists, and data training opted out by default. Streaming PHI redaction is nearing release.

Medical Mode

Compliance

Not a checkbox — a foundation

AssemblyAI offers a Business Associate Agreement (BAA) for organizations subject to HIPAA, included at no additional cost. SOC 2 Type 2, ISO 27001:2022, and PCI DSS v4.0 certified, with data training opted out by default.

For complete data isolation, a self-hosted deployment runs as Docker containers with NVIDIA GPU support — up to 48 concurrent sessions per container.

Certifications & controls

  • BAA available
  • SOC 2 Type 2
  • ISO 27001:2022
  • PCI DSS v4.0
  • Self-hosted option

The paths

Two paths for healthcare deployments

Both run on the same Universal-3.5 Pro Realtime foundation with Medical Mode — build a front-office agent, or an ambient scribe for the exam room.

Recommended

Voice Agent API

Patient intake, triage, scheduling, front-office

$4.50 /hr

Medical Mode works in-pipeline

  • Patient intake and triage voice agents
  • Tool calling for EHR lookups, scheduling, prescription checks
  • Medical Mode terminology accuracy in the pipeline
  • Session resumption keeps context through drops
  • Best for: front-office and patient-facing agents
Get started

Bring your own stack

Universal-3.5 Pro Realtime STT

The ambient scribe architecture

$0.45 /hr

+ $0.15/hr Medical Mode add-on

  • Stream exam-room audio → Medical Mode + diarization + PHI redaction
  • far_field voice_focus for room capture
  • SpeakerRevision for provider/patient/family
  • Pipe structured output to your SOAP-note generator and EHR
  • Best for: ambient clinical documentation
Learn more

Proof

Built for teams who can't get it wrong

Specialized infrastructure lets clinical teams focus on workflow, not a general-purpose speech pipeline.

Healthcare readiness

  • Medical terminology accuracy Pass
  • Speaker separation Pass
  • PHI redaction Pass
  • BAA available Pass
  • Data isolation option Pass
  • JotPsych (behavioral health documentation). “In the medical context, accuracy is highly important, and there can be multiple people present — separating them is key. The biggest impact AssemblyAI has had is enabling our team to focus on workflow features rather than a general speech-to-text pipeline.”

  • Careship. 36% improvement in WER, turning qualitative research into better healthcare experiences across Europe.

  • Works across specialties. Oncology, cardiology, primary care — no retraining.

  • Structured, EHR-ready output. Diarized, PHI-redacted transcripts feed SOAP-note generation.

Build clinical voice AI on infrastructure built for it

Get your API key, or talk to us about a BAA and self-hosted deployment.

Frequently asked questions

What is the best speech-to-text API for medical transcription?

AssemblyAI's Medical Mode posts a 3.2% Missed Entity Rate — the lowest across benchmarked providers (Deepgram, Speechmatics, AWS, Google) — and is purpose-built for medication names, procedures, conditions, and dosages. It activates with a single parameter (domain: "medical-v1") on both pre-recorded and streaming models and works across specialties without retraining.

How do I build an AI medical scribe?

Stream exam-room audio to Universal-3.5 Pro Realtime with Medical Mode, speaker diarization, and PHI redaction enabled, then pipe the diarized, structured transcript into your SOAP-note generator and EHR. voice_focus (far_field) handles ambient room capture and SpeakerRevision disambiguates provider, patient, and family. Pricing is $0.45/hr base plus a $0.15/hr Medical Mode add-on.

Can you build a voice agent for telehealth triage?

Yes. The Voice Agent API handles patient intake and triage end-to-end — one WebSocket for STT, LLM, and TTS at $4.50/hr — with tool calling for EHR lookups and scheduling, Medical Mode terminology accuracy in the pipeline, and session resumption so a dropped call doesn't restart the interaction.

How does AssemblyAI Medical Mode compare to Deepgram Nova-3 Medical?

On medical entity accuracy, Medical Mode posts a 3.2% Missed Entity Rate versus 4.7% for Deepgram Nova-3 Medical — and leads the field against Speechmatics Enhanced Medical (3.6%), AWS Transcribe Medical (8.7%), and Google Medical Conversation (24.4%) as well. See the full benchmark table for methodology.

Does AssemblyAI support HIPAA compliance for voice data?

AssemblyAI is considered a business associate under HIPAA and offers a standard Business Associate Agreement (BAA) — at no additional cost — for customers processing protected health information (PHI). It also provides PHI redaction across audio and transcripts, SOC 2 Type 2, ISO 27001:2022, and PCI DSS v4.0, plus a self-hosted deployment option for complete data isolation.

How much does Medical Mode cost?

Medical Mode is a $0.15/hr add-on on top of the model you're using — $0.45/hr for Universal-3.5 Pro Realtime streaming, or included in the $4.50/hr Voice Agent API pipeline.