Build a voice agent for telehealth triage
Telehealth triage voice agent tutorial: build real-time AI calls with OPQRST symptom capture, severity scoring, care-level routing, and BAA-backed PHI controls



A telehealth triage voice agent is a phone-based agent that answers an inbound patient call, screens for red-flag symptoms, walks a clinical intake protocol, and routes the caller to the right level of care. It does not diagnose and it does not recommend treatment — it collects structured information accurately, scores it against rules a clinician wrote, and escalates to a human. This guide covers the call flow, the speech configuration that makes it safe, the telephony path, and the contracting you need in place before a single real patient calls.
Triage is one of the highest-value structured voice use cases in healthcare, and one of the least forgiving. A scheduling agent that mishears a date costs someone an afternoon. A triage agent that mishears "chest pressure" as "chest congestion" costs something else entirely.
What a telehealth triage voice agent does — and what it must never do
It does not diagnose
Diagnosis is a licensed clinical act. An agent that offers one is practicing medicine, and no amount of prompt engineering makes that acceptable. The agent must never name a condition, suggest a medication, estimate how serious something is, or say anything that could discourage a caller from seeking care.
This is not only a legal constraint. It is what makes the system testable. An agent that only collects and routes has a bounded failure surface: it can mishear, it can mis-score, it can route wrong. Each of those is detectable by a nurse reviewing the record. An agent that reasons toward a conclusion has a failure surface no one can enumerate.
What it does do
- Executes a fixed intake protocol — OPQRST for pain, or whatever your nurse line already uses
- Screens for red-flag symptoms in the first thirty seconds, before anything else
- Captures the chief complaint in the patient's own words, verbatim
- Extracts structured fields: onset, duration, severity, medications, allergies, relevant history
- Scores those fields against deterministic rules written by clinicians
- Routes: emergency transfer, nurse callback, same-day appointment, self-care guidance from an approved script
- Writes a documented record that a human reviews
Where the value sits
The value is not in replacing the nurse. It is in the twenty minutes of structured intake that happen before the nurse gets involved, and in the fact that every call gets the same protocol at 3 a.m. as at 3 p.m. Healthcare AI infrastructure companies like Commure have built this pattern into ambient documentation products: capture the encounter accurately, structure it, hand a clinician something already organized.
The call flow, stage by stage
A triage call has seven stages. The order matters — specifically, red-flag screening comes third, not last.
- Greeting and scope. The agent identifies itself as an automated intake line, states that it does not provide medical advice, and tells the caller how to reach a human immediately.
- Identity verification. Name, date of birth, and one additional identifier, matched against the patient record.
- Red-flag screen. A short fixed battery: chest pain, difficulty breathing, sudden weakness or numbness, uncontrolled bleeding, altered consciousness, suicidal ideation. Any hit exits the protocol immediately.
- Chief complaint capture. Open question, patient's own words, recorded verbatim into the record.
- Protocol questions. The branching intake — onset, provocation, quality, radiation, severity, timing — driven by the complaint category.
- Severity scoring and disposition. Deterministic, in code, from the extracted fields.
- Handoff and documentation. Transfer, callback scheduling, or appointment booking, plus a structured record written for clinical review.
Stages 3 and 4 are where speech accuracy stops being a quality metric and starts being a safety property.
Why the speech layer decides whether this is safe
Everything downstream — the scoring, the routing, the record — runs on the transcript. A transcription error in stage 3 does not produce a slightly worse outcome; it produces a confidently wrong one, because the scoring logic has no way to know the input was wrong.
Clinical entity accuracy: Medical Mode
Medical Mode is a domain setting, not a different model. You add one parameter:
{ "domain": "medical-v1" }It costs +$0.15/hr on top of your base model rate, and it reduces Missed Entity Rate on drugs, conditions, procedures, and clinical terms by roughly 20% compared to running without it. It is available in English, Spanish, German, and French.
To be clear about what that means: Medical Mode is not a prerequisite for clinical accuracy, and Universal-3.6 Pro Realtime handles ordinary medical vocabulary well without it. Medical Mode narrows a specific failure mode — the model dropping or substituting a drug name, a procedure, or a condition it heard correctly but did not weight as clinically salient. On a triage line where "metoprolol" and "metronidazole" have very different implications, narrowing that failure mode is worth $0.15/hr.
Contextual prompting with the patient's prior-visit note
This is the highest-leverage configuration available to a triage agent and it is underused.
Universal-3.5 Pro supports contextual prompting — you prime the model with domain knowledge or prior context before the audio arrives. In an internal healthcare test, feeding a patient's prior-visit note into the prompt cut missed medical terms by 31%.
That result maps almost perfectly onto triage. By the time a patient calls the nurse line, you usually already know their active medication list, their chronic conditions, and what they were last seen for. Passing that in means the model is not guessing at "Eliquis" cold — it is expecting it. Callers on a nurse line are frequently describing a flare of something already in their chart.
The operational requirement is that you fetch the prior-visit note after identity verification (stage 2) and before the protocol questions (stage 5), which is exactly where your patient lookup already sits.
agent_context: making short symptom replies resolve
Triage answers are short. "Since Tuesday." "About a seven." "The left one." "No, the other pill." These are the utterances speech models get wrong most often, because there is almost no acoustic material to work with and no surrounding sentence to disambiguate them.
agent_context fixes this by passing the agent's own question into the transcription request. The model transcribing "about a seven" knows the agent just asked "on a scale of one to ten, how bad is the pain right now?" — and resolves accordingly.
Across a benchmark of 20,000 voice agent audio files, agent_context cut word error rate by 10.2%. The component improvements matter more than the headline for triage: short-utterance errors dropped 13.7%, and medical entities improved 9.4%.
On a triage line, every single turn is a known question. There is no reason not to pass it.
voice_focus and accuracy modes for real patient audio
Patients call from cars, kitchens, waiting rooms, and beds with a television on. Two settings handle this:
- voice_focus isolates the primary speaker and suppresses background noise. Use near-field for callers on a phone handset or headset — which is nearly all inbound triage traffic. Use far-field for rooms, kiosks, and speakerphone-in-the-corner situations. It costs +$0.10/hr.
- Modes replace low-level tuning flags. balanced is the default. Switch to max_accuracy when audio conditions are poor — a noisy home environment, a weak cellular connection, a caller who is short of breath. On a triage line, the small latency cost of max_accuracy is a trade worth making; nobody abandons a nurse-line call over 150 milliseconds.
Turn-taking that doesn't interrupt a patient in distress
End-of-turn detection in Universal-3.6 Pro Realtime evaluates what the patient has actually said, not merely how long they have paused. That matters here more than in most agent use cases: a patient describing chest pain pauses mid-sentence to breathe. A silence-threshold model interprets that as a completed turn and cuts them off. A model that waits for a complete thought keeps listening.
The Voice Agent API tightens this further. On the Sesame TurnBench evaluation it produces 55% fewer false-positive interruptions than Streaming STT alone at the same settings — 3.1 wrong interruptions per conversation versus 6.8 — trading seven points of recall for that. On a triage call, that trade is obviously correct.
Which languages a triage agent can actually run in
This is the constraint most teams discover late, so it is worth stating plainly and separately.
The Voice Agent API supports six languages: English, Spanish, French, German, Italian, and Portuguese, with native code-switching across all six. That is the full conversational surface for an agent built on the managed API.
Medical Mode covers four of those six: English, Spanish, German, and French. Italian and Portuguese triage calls will run, and Universal-3.6 Pro Realtime will transcribe them, but they will not get the clinical-entity domain weighting.
If you need broader coverage on the transcription side specifically, the assembled-stack route gives you more room: Universal-3.5 Pro code-switches natively across 18 languages for async speech-to-text, and Universal-3.6 Pro Realtime covers 32 languages for streaming. Full details are in the supported languages breakdown.
Run a recorded intake call through Universal-3.5 Pro with Medical Mode enabled and see what the transcript looks like before you design a single prompt.
speech_models vs speech_model: which parameter goes where
This trips up almost everyone building a system that does both live and post-call transcription, so here it is explicitly:
| Parameter | Product | Shape | Example |
|---|---|---|---|
speech_models (plural) |
Pre-recorded / async transcription | List | speech_models=["universal-3-5-pro"] |
speech_model (singular) |
Streaming / real-time transcription | String | speech_model="universal-3-6-pro" |
The model IDs differ too: universal-3-6-pro for streaming, universal-3-5-pro for async. A triage system uses both: singular in the live WebSocket session, plural in the post-call async pass that produces the record of truth.
Two ways to build it
| Dimension | Voice Agent API | Assembled stack |
|---|---|---|
| Integration | One WebSocket covering STT, LLM, and TTS | Three services you wire together and keep in sync |
| Pricing | Flat $4.50/hr ($0.075/min), one bill | $0.45/hr streaming base, plus add-ons, plus LLM and TTS |
| Languages | Six, with native code-switching | 32 for streaming speech-to-text (18 async); LLM and TTS coverage is yours to source |
| Latency tuning | One place; ~1s end-to-end, under 500ms over SIP | Yours to optimize across three network hops |
| Model chain control | Managed loop; less component-level control | Full — swap any component independently |
| Healthcare contracting | Contact sales to scope a healthcare deployment | BAA-backed deployments run through Speech-to-Text and Streaming STT |
| Best fit | Getting a safe, bounded agent live quickly with predictable cost | Existing orchestration, unusual requirements, or self-hosted deployment |
The contracting row is the one that usually decides it for healthcare teams. Read it before you pick.
Getting the triage line on a phone number
A triage agent is a phone product. Patients call a number; they do not open a web app.
The Voice Agent API supports SIP telephony — it plugs into any SIP endpoint, via Twilio, an AssemblyAI-owned number, or a number you already own and are already routing through your call center. Latency over the SIP path is under 500ms. DTMF is supported too, which matters more in healthcare than most verticals: keypad entry is often the right input for a date of birth or a member ID, and it avoids putting an identifier through speech recognition at all.
Webhooks cover both session and telephone-call lifecycle events, which is what you hook your audit logging into. Connection details are in the Twilio integration docs.
Two practical notes for triage specifically:
- Warm transfer, not blind transfer. When the red-flag screen fires, the agent needs to hand the nurse the transcript and extracted fields along with the call. Configure the escalation tool to write the record before it initiates the transfer.
- Session resumption. Reconnecting within 30 seconds preserves context. Patients on cell phones drop calls; a dropped call mid-protocol should not restart the protocol.
Building the transcription layer
Two passes. The live pass drives the conversation. The async pass produces the record a clinician and a regulator will actually read.
Live: streaming configuration
import asyncio
import json
import websockets
WS_URL = "wss://streaming.assemblyai.com/v3/ws"
async def open_triage_stream(current_agent_question: str, prior_visit_note: str):
async with websockets.connect(
WS_URL,
additional_headers={"Authorization": "YOUR_API_KEY"},
) as ws:
await ws.send(json.dumps({
"type": "configure",
# SINGULAR: speech_model is the streaming parameter
"speech_model": "universal-3-6-pro",
"domain": "medical-v1", # Medical Mode, +$0.15/hr
"sample_rate": 8000, # PSTN telephony audio
"voice_focus": "near-field", # caller on a handset or headset
"mode": "max_accuracy", # noisy home environments
# Pass the agent's own question so short replies resolve
"agent_context": current_agent_question,
# Prime with the patient's prior-visit note
"prompt": prior_visit_note,
}))
async for message in ws:
event = json.loads(message)
if event.get("type") == "Turn" and event.get("end_of_turn"):
yield event["transcript"]
Three of those settings are the ones that move the needle on a triage call: domain for clinical vocabulary, agent_context for the short answers, and prompt for the patient's history. sample_rate: 8000 is not optional — telephony audio is 8 kHz and telling the model otherwise degrades everything.
Update agent_context on every turn. It is the current question, not a static string.
Post-call: the async record of truth
import assemblyai as aai
aai.settings.api_key = "YOUR_API_KEY"
config = aai.TranscriptionConfig(
# PLURAL: speech_models is the pre-recorded parameter
speech_models=["universal-3-5-pro"],
speaker_labels=True,
redact_pii=True,
redact_pii_audio=True,
redact_pii_policies=[
aai.PIIRedactionPolicy.person_name,
aai.PIIRedactionPolicy.date_of_birth,
aai.PIIRedactionPolicy.phone_number,
],
)
transcript = aai.Transcriber().transcribe("triage_call.wav", config=config)
Run this against the full call recording after the agent hangs up. It has the whole audio file rather than a stream, so it produces a cleaner transcript than the live pass, and it gives you speaker labels and PII redaction across both the transcript and the audio itself. This is the artifact that goes into the chart. The live transcript is a control signal; this is the record.
<div class="blog-cta_component">
<div class="blog-cta_title">Try the streaming model on triage audio</div>
<div class="blog-cta_rt w-richtext">
<p>Drop a sample intake recording into the playground, toggle Medical Mode on and off, and compare how the clinical terms come through.</p>
</div>
<a href="https://www.assemblyai.com/playground" class="button w-button">Try playground</a>
</div>
The parts that make it safe
Speech accuracy is necessary and nowhere near sufficient. Six things do the rest.
A system prompt with hard refusals
Explicit, enumerated, non-negotiable: never name a condition, never suggest or adjust a medication, never estimate how serious something is, never discourage the caller from seeking care, never speculate about causes. Then test the refusals under pressure — patients ask "do you think it's a heart attack?" directly, and often more than once.
Tool calling with a narrow surface
Keep the tool set small and boring: patient lookup, record symptom, compute severity score, escalate to human, schedule appointment, write documentation. Every tool is idempotent and every call is logged. A small surface is an auditable surface.
Tool-call accuracy on AssemblyAI's internal dataset improved from 69% to 80%, and with tool parameter guardrails in place, hallucinated tool arguments went from 2.5% of 3,000 calls to 0% — arguments can only be inferred from user turns and tool results, never from the model's own generations. That last property is the one that matters clinically. A model cannot invent a severity value that no one said. Details are in the tool calling docs.
Deterministic severity scoring
The model extracts fields. Code scores them. Use the scoring logic your nurse line already runs, unchanged.
This is the single most important architectural decision in the system. Scoring in code means the scoring is reviewable by a clinician who does not read Python, testable against historical calls, versionable, and identical on every call. Scoring in the model means none of those things.
Escalation that fails toward a human
Every ambiguous path ends at a person. Low transcription confidence, audible distress, a tool failure, a caller who goes off-protocol, a caller who asks for a human — all of them transfer. The design goal is not to minimize transfers. It is to make sure no caller who needed a human failed to get one.
Audit logging built for a regulator
Log, per call: the audio reference, the full transcript, extracted entities with timestamps, every tool call with arguments and results, the computed severity score with its inputs, the disposition, and the model and configuration version in effect. Version the config. When someone asks in eighteen months why a call was routed the way it was, "we changed the prompt sometime that quarter" is not an answer.
Testing before any patient hears it
Build a test suite before you build the agent. It should cover every red-flag phrase and its common colloquial variants, every protocol branch, accented speech, low-quality phone audio, and background noise. Then pilot with a nurse listening to every call live, with the ability to break in. Not a sample. Every call, until the transfer rate and the disagreement rate both stabilize.
PHI, consent, and contracting
Every triage call is protected health information from the greeting onward.
Disclose at the top of the call. The caller should know they are talking to an automated intake system, that the call is recorded, and how to reach a human immediately. Say it in the first ten seconds, before the identity check.
On HIPAA: AssemblyAI enables covered entities and their business associates subject to HIPAA to use the AssemblyAI services to process protected health information (PHI). AssemblyAI is considered a business associate under HIPAA, and we offer a standard Business Associate Addendum (BAA) that is required under HIPAA to ensure that AAI appropriately safeguards PHI. You can read the terms on the BAA legal page or in the BAA FAQ.
Scoping caveat, and it matters for this post: BAA-backed deployments run through the Speech-to-Text and Streaming STT products. If you are building a triage line on the Voice Agent API specifically, contact sales to scope it — do not assume the managed agent path is covered by the same arrangement.
The rest of the compliance surface: PII and PHI redaction is available across both audio and transcripts. AssemblyAI holds SOC 2 Type 2, ISO 27001:2022, and PCI DSS v4.0. Self-hosted deployment and EU data residency (api.eu.assemblyai.com, streaming.eu.assemblyai.com, same price as US) are available where data residency requirements apply.
Get the BAA in place before the pilot, not before launch. Pilot calls are real PHI from the first one.
Do not deploy a triage agent without all four of these:
- A signed BAA covering the services you are actually using
- Clinical review and sign-off on the protocol, the refusal list, and the scoring logic
- A human escalation path that is always available and always reachable
- IRB or institutional compliance review, depending on your setting
Any one of these missing means you are not ready, regardless of how good the transcripts look.
What a triage line costs to run
Cost is usually the second question after safety, so here is the arithmetic on both paths.
Voice Agent API: flat $4.50/hr ($0.075/min), bundling speech-to-text, LLM, and text-to-speech through one WebSocket. Keyterms prompting is included at no extra charge. One bill, one set of logs. A ten-minute triage call costs $0.75.
Assembled stack, live leg: $0.45/hr for Universal-3.6 Pro Realtime, plus $0.15/hr for Medical Mode, plus $0.10/hr for voice focus — call it $0.70/hr for the transcription layer, before your LLM and TTS costs. Keyterm prompting is included on Universal-3.6 Pro Realtime.
Post-call async pass: $0.21/hr for Universal-3.5 Pro, plus $0.15/hr for Medical Mode, plus $0.02/hr for diarization, plus $0.08/hr for PII text redaction and $0.05/hr for PII audio redaction. That is $0.51/hr for a fully redacted, speaker-labeled clinical record — and calls are short, so this is cents per encounter.
Current rates for every add-on are on the pricing page.
Where triage agents go next
The direction is more context, not more autonomy. The agent should not get smarter about medicine; it should get better informed about the specific patient.
Contextual prompting with the prior-visit note is the first step, and the 31% reduction in missed medical terms is a strong signal that it works. The next steps are the same idea extended: active medication list, known allergies, recent lab flags, the disposition of the last three calls. All of it primes transcription, none of it changes the agent's authority. The agent still collects and routes. It just stops being surprised by the vocabulary.
The second direction is coverage. Triage lines are busiest at night and on weekends, which is exactly when nurse staffing is thinnest. An agent that handles structured intake for the queue — accurately, identically, every time — is the version of this that clinicians actually want.
Start with the transcript. If the transcript is wrong, nothing built on top of it can be right.
Healthcare deployments need a BAA and the right product scope before a single pilot call. Talk to us about what your triage line needs.
Frequently asked questions
How do I build a telehealth triage voice agent?
Build it in seven stages: greeting and scope disclosure, identity verification, red-flag screening, chief complaint capture, branching protocol questions, deterministic severity scoring, and handoff with documentation. Run the red-flag screen third — within the first thirty seconds — not at the end. Extract structured fields with a speech model configured for clinical audio, score those fields in code using your existing nurse-line logic rather than in the model, and route every ambiguous case to a human. Get a signed BAA, clinical sign-off, and compliance review in place before the pilot, because pilot calls carry real PHI.
Can a telehealth voice agent diagnose patients?
No. Diagnosis is a licensed clinical act, and a voice agent must never name a condition, recommend or adjust a medication, estimate how serious a symptom is, or say anything that could discourage a caller from seeking care. The agent's job is to collect structured information accurately and route the caller to the right level of care. This constraint is also what makes the system safe to operate: an agent that only collects and routes has a bounded, testable failure surface, while an agent that reasons toward a clinical conclusion does not.
How many languages does the Voice Agent API support?
The Voice Agent API supports six languages — English, Spanish, French, German, Italian, and Portuguese — with native code-switching across all of them. Medical Mode, the clinical-vocabulary domain setting, covers four of those six: English, Spanish, German, and French. If you need broader transcription coverage, Universal-3.5 Pro code-switches natively across 18 languages for pre-recorded speech-to-text, and Universal-3.6 Pro Realtime covers 32 languages for streaming, on the assembled-stack path.
What is Medical Mode and do I need it for a triage agent?
Medical Mode is a domain setting you enable by passing "domain": "medical-v1" — no model switch required. It costs an additional $0.15/hr and reduces Missed Entity Rate on drugs, conditions, procedures, and clinical terms by roughly 20%, in English, Spanish, German, and French. It is not required for clinical accuracy; Universal-3.5 Pro handles general medical vocabulary well without it. It narrows a specific failure mode — dropped or substituted drug names, procedures, and conditions — which is worth the cost on a line where those terms drive routing decisions.
What is the difference between speech_models and speech_model?
They are the same setting for two different products. Plural speech_models=["universal-3-5-pro"] is the pre-recorded (async) transcription parameter and takes a list. Singular speech_model="universal-3-6-pro" is the streaming parameter and takes a string. A triage system typically uses both: the singular form in the live WebSocket session that drives the conversation, and the plural form in the post-call async pass that produces the clinical record.
How does the agent handle short or mumbled answers to triage questions?
Pass the agent's own question into the transcription request using agent_context. Triage replies are short — "since Tuesday," "about a seven," "the left one" — and short utterances have the least acoustic information to work with, so they are where speech models fail most. Across a benchmark of 20,000 voice agent audio files, agent_context cut word error rate by 10.2%, with short-utterance errors down 13.7% and medical entities improved 9.4%. On a triage line every turn is a known question, so there is no reason not to pass it.
What happens when a caller reports a red-flag symptom?
The agent exits the protocol immediately and does not continue asking intake questions. Screen for red flags — chest pain, difficulty breathing, sudden weakness or numbness, uncontrolled bleeding, altered consciousness, suicidal ideation — in the first thirty seconds of the call, not at the end. On a hit, the agent reads a fixed escalation script written and approved by clinicians, writes the record, and initiates a warm transfer so the receiving nurse gets the transcript and extracted fields along with the call.
Is AssemblyAI HIPAA-compliant, and how does it handle PHI?
AssemblyAI enables covered entities and their business associates subject to HIPAA to use the AssemblyAI services to process protected health information (PHI). AssemblyAI is considered a business associate under HIPAA, and we offer a standard Business Associate Addendum (BAA) that is required under HIPAA to ensure that AAI appropriately safeguards PHI. BAA-backed deployments run through the Speech-to-Text and Streaming STT products — for Voice Agent API healthcare use cases, contact sales to scope the deployment. PHI redaction is available across audio and transcripts, and AssemblyAI holds SOC 2 Type 2, ISO 27001:2022, and PCI DSS v4.0, with self-hosted deployment and EU data residency available where residency requirements apply.
Lorem ipsum dolor sit amet, consectetur adipiscing elit, sed do eiusmod tempor incididunt ut labore et dolore magna aliqua. Ut enim ad minim veniam, quis nostrud exercitation ullamco laboris nisi ut aliquip ex ea commodo consequat. Duis aute irure dolor in reprehenderit in voluptate velit esse cillum dolore eu fugiat nulla pariatur.




