What is speech understanding?
Speech understanding turns audio into structured insight: sentiment, entities, topics, moderation, and PII redaction. What it is and when to use it.



A transcript is not an answer. It's a wall of text that happens to be shaped like a conversation, and if you've ever tried to build a product on top of raw transcripts you know the feeling: you have all the words and none of the meaning.
Speech understanding is the layer that closes that gap. It takes the output of transcription — words, timestamps, speakers — and produces structured fields you can actually query: who felt what, which entities were named, what the conversation was about, whether anything in it violates policy, and which sixty seconds a human should listen to. (This capability was previously marketed as "audio intelligence," which you may still see in older docs and integrations.)
I run Voice AI at AssemblyAI, and the thing I'd most like to correct about how this category gets explained is the implication that it's a nice add-on. For most products built on voice, it's the actual product. Transcription is the cost of entry; what you do with the transcript is the thing customers pay for.
So: what speech understanding is, what's in it, how each piece works, and where the real engineering difficulty sits.
Speech understanding, defined
Speech understanding is the set of AI capabilities that extract structured meaning from spoken audio — sentiment, entities, topics, key moments, content safety signals, and redacted personal information — on top of a transcript.
Three properties distinguish it from transcription:
It produces schema, not prose. A transcript is a string. Speech understanding output is typed: a sentiment score attached to a speaker and a time range, an entity with a category and a character offset, a topic with a relevance weight. Typed output is queryable; prose isn't.
It's per-segment, not per-file. A single sentiment label for a 40-minute call is nearly useless. The signal is in the trajectory — where the conversation turned, and who turned it.
It's grounded in the audio, not just the text. The good implementations carry timestamps and speaker attribution all the way through, so every extracted fact points back at the moment it came from. That's what makes the output auditable rather than merely plausible.
You'll also see this described as speech intelligence or, when applied specifically to customer and sales conversations, conversation intelligence. Same underlying capability set, different packaging depending on who's buying.
The full feature set
Here's what's actually in AssemblyAI's Speech Understanding API, organized by what each one is for rather than by how it's implemented.
| Feature | What it returns | What it's for |
|---|---|---|
| Entity detection (50+ types) | Typed entities with positions | CRM enrichment, search, structured extraction |
| Sentiment analysis | Per-segment sentiment | Escalation detection, CSAT proxies, coaching |
| Topic detection (IAB) | Topic labels with relevance scores | Contact-driver reports, brand safety, routing |
| Key Phrases | Ranked significant phrases (English) | Jump-to moments, highlight reels |
| Content safety / moderation | Policy labels, scores, timestamps | Review queues, compliance flags |
| PII redaction (text + audio) | Redacted transcript and audio file | Privacy compliance before storage |
| Profanity filtering | Masked transcript text | Family-safe output, QA flags |
| Translation (100+ targets) | Translated transcript (pre-recorded only) | Multilingual review and distribution |
| Audio event tagging (100+ tags) | Non-speech event labels | Context that isn't spoken — alarms, music, laughter |
| Automatic language detection | Detected language (130+ detectable) | Routing multilingual pipelines (needs 15+ seconds) |
| Custom spelling & formatting | Normalized transcript output | Brand and product names rendered correctly |
| Multichannel | Per-channel transcripts | Stereo call recordings with separated legs |
Anything that doesn't fit a fixed schema — arbitrary questions, custom summaries, structured extraction against your own JSON shape — runs through the LLM Gateway, a single API to OpenAI, Anthropic, Google, and other providers, pointed at your transcripts. That's the escape hatch, and it's the right one: fixed features for the things you'll run on every file, a general model for the things that change with your product.
Two naming notes, because both come up in older integrations. Key Phrases is the current name for what used to be branded "Automatic Transcript Highlights"; the feature is live, the old label isn't. And Auto Chapters is genuinely deprecated — if you're still calling auto_chapters, the migration path is Summarization under Speech Understanding.
Get sentiment, entities, topics, and redaction from a single transcription request. Free account, no sales call, docs you can read in an afternoon.
What are Voice AI Guardrails, and which work with pre-recorded speech-to-text?
Guardrails are the safety, compliance, and quality controls that run alongside transcription rather than after it — the layer that keeps a voice product from saying, storing, or surfacing something it shouldn't.
They divide into two groups by what they protect.
Content and compliance guardrails operate on what was said: content safety classification against a policy taxonomy, profanity detection and masking, and PII redaction across both the transcript and the audio file itself. The audio side is the one teams forget. A transcript with a card number masked is not much protection if the recording still reads it aloud — redacting both is what actually shrinks the scope of a security review.
Quality and cost guardrails operate on the system: confidence signals you can threshold on, so low-certainty output routes to review rather than straight into a customer-facing summary, and controls that keep a runaway session from becoming a runaway bill. There's a fuller treatment in built-in protection for compliance, quality, and cost control.
For pre-recorded speech-to-text specifically, these run in the same async request as transcription: content safety, profanity filtering, PII redaction in text and audio, entity detection, topic detection, sentiment analysis, and audio event tagging. You don't call a second service — you set the parameters on the transcription request and the structured output comes back alongside the transcript, sharing its timestamps and speaker labels.
That shared-timestamp property is worth dwelling on. When a content safety flag, a sentiment dip, and a speaker turn all reference the same time index, you can reconstruct exactly what happened. When they come from three vendors that each segmented the audio differently, you can't — and reconciling them becomes its own project.
For live audio, the streaming path at wss://streaming.assemblyai.com/v3/ws supports classifying partial transcripts as they arrive. More on that below.
How does speaker separation work?
Speaker separation — diarization — answers "who spoke when," and it's the feature most likely to quietly ruin everything downstream when it's wrong.
Here's the mechanism in the current flagship. Rather than transcribing first and clustering speakers afterward as a separate post-process, Universal-3.5 Pro jointly produces the transcript and the speaker change points. The model knows where a turn ended because it's the same model that decided what the words were. That matters most in exactly the conditions where naive diarization collapses: short turns ("mm-hm," "right," "no wait"), rapid back-and-forth, and overlapped speech where two people talk at once.
It's also optimized for the right metric. Most diarization is evaluated with DER (diarization error rate), which measures how much audio time got assigned to the wrong speaker. That's a fine acoustic metric and a poor product metric, because it doesn't care whether the words ended up attributed correctly. Universal-3.5 Pro is optimized for cpWER — concatenated minimum-permutation word error rate — which measures the thing you actually consume: did each speaker's words land under that speaker? Average cpWER is 30.17, against 37.92 for Deepgram Nova-3 English and 35.26 for ElevenLabs Scribe v2.
For live audio, streaming diarization works differently and cleverly: it labels speakers as the stream runs, then re-clusters once it has heard the whole conversation and sends a single correction within roughly half a second of the stream ending. Up to 10 speakers. You get usable labels in real time without being permanently stuck with the early, under-informed guess. Streaming speaker diarization goes into the mechanics.
One configuration note that catches people: turning on speaker labels changes streaming's turn-detection behaviour, shifting the silence thresholds and disabling continuous partials. If your partials stop flowing the day you enable diarization, that's why.
Practical guidance that saves people weeks: if your audio is stereo with the participants already on separate channels — most call recordings are — use multichannel transcription instead of diarization. Channel separation is ground truth; diarization is inference. Don't infer something you were handed.
Can conversations be analyzed in real time?
Yes, and the architecture is genuinely different from batch analysis rather than just faster.
The streaming path uses Universal-3.5 Pro Realtime over the v3 WebSocket, returning partial and final transcripts within a few hundred milliseconds. Speech understanding then runs over those partials, which means sentiment, entities, and content safety signals arrive while the conversation is still happening — early enough to change it.
Three streaming-specific capabilities are worth knowing about, because they don't exist in the batch world:
agent_context. You pass in what your agent just asked, and the model uses it to resolve exactly the utterances that are hardest without context — a spelled-out email address, a one-word confirmation, a mumbled account number. Across a benchmark of more than 10,000 voice agent audio files, passing agent context cut word error rate by 8.9%, rising to 16.4% when a context prompt is supplied alongside it, with place-name entity errors down 30.7% and fabrications down 27.0%. It's included in the base rate, so there's no reason not to use it.
Context Carryover. A short rolling conversation memory, on by default, so the model carries what was established earlier in the call into how it interprets what comes later. David Zhao, Co-founder at LiveKit, singled this out:
“We're excited to make AssemblyAI's Universal-3.5 Pro available on LiveKit Inference. What really stands out is their pace of innovation with Context Carryover — it intelligently applies conversation context to improve transcription accuracy in a way most speech models don't, removing the need for users to predefine key terms.”
voice_focus. Isolates the primary speaker and suppresses background speech and noise, with near-field for headsets and phones and far-field for rooms, kiosks, and drive-thrus. Note the hyphens — those are the only accepted values, and an underscored variant is rejected outright with a validation error that closes the session. Real-time analysis in a noisy environment is mostly a signal-isolation problem before it's a modeling problem.
There are also three turn-detection modes instead of a pile of low-level flags — min_latency, balanced (the default), and max_accuracy for noisy or far-field audio. Rather than simply waiting out a silence timer, the model evaluates whether what has been said is complete, and combines that judgement with the mode's silence thresholds (min_turn_silence 128/128/512 ms and max_turn_silence 640/1280/2560 ms across the three modes, with vad_threshold defaulting to 0.2). That's why reading out an account number digit by digit doesn't get chopped in half the way a pure silence timer chops it.
The design question for real-time isn't "can I get the transcript fast." It's "what do I do differently because I have it now?" Real-time conversation intelligence covers the patterns that earn their keep: live agent assist, in-call compliance prompts, and escalation detection before the customer asks for a supervisor. The product surface is on the streaming speech-to-text page.
Stream your own audio and watch partial transcripts, speaker labels, and extracted insight appear as you speak.
How do you surface the significant moments in a recording?
Nobody listens to the whole file. The product job is to find the 90 seconds that matter and put the user there.
Three mechanisms do most of the work, and they're complementary rather than competing.
Key Phrases ranks the phrases in a transcript by significance and returns them with timestamps. This is the cheapest useful thing you can run on a recording and it's often enough on its own for a jump-to interface. It's English-only.
Sentiment trajectory. Per-segment sentiment is a curve, not a number, and the steepest slope is usually the moment worth reviewing. A call that runs neutral for 18 minutes and drops hard at minute 19 has told you exactly where to look — and, with speaker attribution attached, told you who caused it.
Entity clustering. When the same entity — a competitor name, a specific policy number, a product SKU — appears repeatedly in a short window, that window is about something. This is the least-used technique on the list and one of the most effective, because it finds significant moments without needing to know in advance what "significant" means for your domain.
For anything domain-specific — "find the moment the customer stated their reason for cancelling," "pull the section where the candidate described the system they built" — send the transcript with its timestamps through the LLM Gateway and ask for time ranges back. The fixed features handle the general case cheaply; the general model handles the case that's specific to your product.
Metaview does this for recruiting conversations, and Shahriar Tajbakhsh, their Co-founder and CTO, described what changed the quality of the output:
“Since moving to AssemblyAI, we've seen a meaningful improvement in the confidence tail of our production transcripts….What stands out is not just the model quality, but the way [they] let us bring real meeting context into transcription, from calendar titles to organizations, domains, and participant names, so recruiting conversations come through with the nuance our customers depend on.”
The phrase to notice is "confidence tail." Average transcript quality is rarely the constraint on a speech understanding product. The tail is — the small share of segments where the model was unsure, which is disproportionately where the names, numbers, and decisions live.
That's what contextual prompting on Universal-3.5 Pro addresses. You prime the model with a domain or prior context — meeting agendas, participant names, product and competitor vocabulary, prior-visit clinical notes — before it transcribes. The Universal-3.5 Pro release post covers how to use it.
How do you identify key topics across hours of audio?
Single-file topic detection is straightforward. Topic analysis across a corpus — a quarter of support calls, a year of podcast episodes, an entire research study — is a different job, and it's where teams most often build something that doesn't survive contact with real data.
The approach that works:
Start with IAB topic detection as your stable spine. A standardized taxonomy gives you labels that mean the same thing in January and in June, which is the whole point of trend analysis. Custom taxonomies drift as whoever maintains them changes their mind.
Layer entity frequency on top. Topics tell you the category; entities tell you the specifics. "Billing" as a topic is not actionable. "Billing" plus a spike in mentions of a specific plan name after a pricing change is a root cause.
Use the LLM Gateway for emergent themes, then promote the durable ones. Run open-ended theme extraction over a sample, see what recurs, and once a theme proves stable, encode it as a keyterm list or a classification rule so it's measured consistently instead of rediscovered every run.
Weight by relevance and duration, not by count. A topic mentioned once in 400 calls and a topic that occupies half of 40 calls both produce a count. Only one is a signal.
One accuracy note that has outsized impact at corpus scale: if the transcription layer inconsistently renders your product names, topic and entity aggregation fragments across spellings, and your trend lines become meaningless. Custom spelling and keyterm prompting fix this at the source. It's an unglamorous configuration step that determines whether corpus analysis works at all — though keep the keyterm list scoped, because an over-broad list can pull the model off words it was already getting right.
Jiminny builds conversation intelligence for revenue teams on top of this kind of pipeline, and Tom Lavery, their CEO and Founder, framed the partnership side of it:
“AssemblyAI has a real high-touch personal service. It's a great partnership and we're very collaborative and get to test new AI models early and work together. And AssemblyAI is really pushing boundaries, helping us create a well-rounded conversation intelligence platform.”
Can it handle enterprise-scale processing?
This is the question that decides deployments, and it has four parts.
Capacity and rate limits. Worth being precise, because "unlimited" is the wrong word and the real mechanism is genuinely favourable. On streaming there is no cap on how many sessions you can hold open at once; what's governed is how many new sessions you open per minute (100+ on paid accounts), and that allowance auto-scales by 10% every minute you run at 70% or more of it, with no ceiling — so sustained load lifts your limit rather than throttling it. Go over the current allowance and new connections are refused with close code 1008 while existing calls continue untouched; the allowance also relaxes back down after quiet periods, so a cold spike is the case to plan for. On pre-recorded, the limit is jobs processing in parallel (200+ on paid) and anything beyond it is queued FIFO rather than rejected. Higher limits on both surfaces are available on request at no additional cost. This matters more than throughput benchmarks suggest, because enterprise audio volume isn't smooth — it's Monday morning, it's the day after an outage, it's the hour a campaign lands.
Deployment shape. Three options, and organizations pick by constraint rather than preference. Voice AI Cloud is AssemblyAI-hosted and the default. EU data residency runs the same models through api.eu.assemblyai.com and streaming.eu.assemblyai.com at the same price, with data staying in the EU. Self-hosted runs inside your own infrastructure — a customer AWS VPC, for instance — for teams with contractual or regulatory limits on where audio can travel.
Cost predictability. Billing is per second with no minimums and no upfront commitments, which means volume modeling is arithmetic rather than negotiation. Current rates for models, add-ons, and speech understanding features are on the pricing page.
Model stability. Pin a specific model when you need reproducible output across a long-running analysis, or omit the model parameter to auto-upgrade to the latest Universal Pro release. Both are legitimate; the mistake is not deciding, and discovering during a compliance review that your Q1 and Q3 numbers came from different models.
For regulated industries, the supporting pieces are PII redaction across audio and transcripts, SOC 2 Type 2, and — for healthcare workloads — Medical Mode via a single domain: "medical-v1" parameter on Universal-3.5 Pro or Universal-3.5 Pro Realtime, plus a standard Business Associate Addendum for customers processing PHI.
Where the difficulty actually is
If you take one thing from this, take this: speech understanding output is only as good as the transcript underneath it, and the failure mode isn't obvious.
A classifier reading a corrupted transcript doesn't return an error. It returns a confident, well-formatted, completely wrong label — and it does that at scale, on every file, in a JSON shape that looks exactly like the correct output. There's no exception to catch. The only way to find it is to measure the transcript separately from everything built on it, which almost nobody does until an incident forces it.
That's the argument for caring about the model underneath more than the feature list on top. The transcription layer determines your ceiling; the speech understanding layer determines how close you get to it. Teams that pick the second without examining the first spend the following year tuning prompts against a problem that lives one layer down.
The other thing worth saying, because it's changed recently: speech understanding used to be a batch category. You processed a recording, you got fields back, you built a dashboard. The interesting products now run the same extraction on a live stream and do something with the result inside the conversation — and that reframes the whole category from reporting to intervention. It's a much better place for this technology to be, and it's a much harder engineering problem, which is usually a good sign about where the value is.
Transcription, sentiment, entities, topics, moderation, and redaction from one API. Pay-as-you-go, auto-scaling limits, EU residency available.
Frequently asked questions
What is speech understanding?
Speech understanding is the set of AI capabilities that extract structured meaning from spoken audio on top of a transcript — sentiment, entity detection, topic classification, key phrases, content safety labels, and PII redaction. Unlike transcription, which returns text, speech understanding returns typed fields with timestamps and speaker attribution, so the output can be queried, charted, and audited. This capability was previously marketed as "audio intelligence."
What's the difference between speech-to-text and speech understanding?
Speech-to-text converts audio into words, timestamps, and speaker labels. Speech understanding takes that output and produces meaning: who felt what, which entities were named, what the conversation was about, and which moments need attention. They run together in a single API request, and the second is capped by the accuracy of the first — a classifier reading a flawed transcript returns a confident wrong answer rather than an error.
Can speech understanding analyze conversations in real time?
Yes. Running over the streaming WebSocket with Universal-3.5 Pro Realtime, partial and final transcripts return within a few hundred milliseconds, and sentiment, entity, and content safety analysis can run on those partials while the conversation is still happening. Streaming adds capabilities batch processing doesn't have, including agent_context — passing in the question your agent just asked, which cut word error rate by 8.9% across more than 10,000 voice agent audio files, or 16.4% with a context prompt alongside it.
How does speaker diarization identify who is speaking?
Universal-3.5 Pro produces the transcript and the speaker change points together rather than clustering speakers in a separate post-process, which is why it holds up on short turns, rapid back-and-forth, and overlapped speech. It's optimized for cpWER — whether each speaker's words landed under the right speaker — at an average of 30.17, versus 37.92 for Deepgram Nova-3 English. For streaming audio, speakers are labeled live and re-clustered with a single correction sent within about half a second of the stream ending, for up to 10 speakers.
What can you use speech understanding for?
The common production uses are conversation intelligence for sales and support calls, meeting and notetaking products that surface decisions and action items, media workflows that generate chapters and highlight reels, compliance systems that flag policy violations with timestamps, and voice agents that adapt mid-call based on detected sentiment or intent. The unifying pattern is turning unstructured speech into fields a product can act on.
Does speech understanding work at enterprise scale?
Yes, and the mechanism is worth knowing rather than taking on faith. Streaming puts no cap on concurrently open sessions and governs new sessions per minute instead (100+ on paid accounts), auto-scaling that allowance 10% per minute at sustained load with no ceiling; pre-recorded caps parallel jobs (200+ on paid) and queues the overflow FIFO rather than rejecting it. Higher limits on either surface are free on request. Billing is per second with no minimums, and deployment options are hosted cloud, EU data residency at the same price, or self-hosted inside your own infrastructure. Supporting controls for regulated workloads include PII redaction across both audio and transcripts, SOC 2 Type 2, Medical Mode for clinical audio, and a standard Business Associate Addendum for customers processing PHI.
Lorem ipsum dolor sit amet, consectetur adipiscing elit, sed do eiusmod tempor incididunt ut labore et dolore magna aliqua. Ut enim ad minim veniam, quis nostrud exercitation ullamco laboris nisi ut aliquip ex ea commodo consequat. Duis aute irure dolor in reprehenderit in voluptate velit esse cillum dolore eu fugiat nulla pariatur.





