New Universal-3.6 Pro Realtime is now available Learn more

Speech Understanding API

Turn raw audio into structured, actionable intelligence. Purpose-built models extract meaning from speech in a single API call.

Features

Output

Lauren: Thank you for calling Nissan. My name is Lauren. Can I have your name?

John: Yeah, my name is John Smith.

Lauren: Thank you, John. How can I help you?

John: I was just calling about to see how much it would cost to update the map in my car.

Lauren: I’d be happy to help you with that today. Did you receive a mailer from us?

John: I did.

Lauren: Do you need the customer number?

John: Yes, please.

Lauren: Okay. It’s 15243. Thank you. And the year, make, and model of your vehicle?

John: Yeah, I have a 2009 Nissan Altima.

Lauren: Oh, nice car.

John: Yeah, thank you. We really enjoy it.

Lauren: Okay, I think I found your profile here. Can I have you verify your address and phone number, please?

John: Yes, it’s 1255 North Research Way. That’s in Orem, Utah 84097. And my phone number is 801-431-1000.

Lauren: Thanks, John. I located your information. The newest version we have available for your vehicle is version 7.6, which was released in March of 2012. The price of the new map is $99 plus shipping and tax. Let me go ahead and set up this order for you.

John: Um, well, can we wait just a second? I’m not really sure if I can afford it right now.

Lauren: Alright, well, here are a few reasons to consider purchasing today. It looks as though you haven’t updated your vehicle for 3 years, so that would be the equivalent of getting 3 years’ worth of updates for the price of 1.

John: Oh, okay.

Lauren: In addition, special offers like the current promotion don’t come around too often. I would definitely recommend taking advantage of the extra $50 off before it expires.

John: Yeah, that does sound pretty good.

Lauren: If I set this order up for you now, it’ll ship out today and for $50 less. Do you have your credit card handy and I can place this order for you now?

John: Yeah, let’s go ahead and use a Visa.

Code

import assemblyai as aai

aai.settings.api_key = "YOUR_API_KEY"

audio_file = "https://assembly.ai/wildfires.mp3"

config = aai.TranscriptionConfig(
    speech_models=["universal-3-5-pro", "universal-2"],
    language_detection=True,
    speaker_labels=True,
)

transcript = aai.Transcriber().transcribe(audio_file, config=config)

if transcript.status == aai.TranscriptStatus.error:
    raise RuntimeError(f"Transcription failed: {transcript.error}")

print(transcript.text)

for utterance in transcript.utterances:
    print(f"Speaker {utterance.speaker}: {utterance.text}")
Metaview
Dovetail
Granola
Apollo.io
Ashby
Siro
Calabrio
Cluely
Genio
Commure
Retell
CallRail
LiveKit
EliseAI
ClickUp
HeyGen
Metaview
Dovetail
Granola
Apollo.io
Ashby
Siro
Calabrio
Cluely
Genio
Commure
Retell
CallRail
LiveKit
EliseAI
ClickUp
HeyGen
Metaview
Dovetail
Granola
Apollo.io
Ashby
Siro
Calabrio
Cluely
Genio
Commure
Retell
CallRail
LiveKit
EliseAI
ClickUp
HeyGen
Metaview
Dovetail
Granola
Apollo.io
Ashby
Siro
Calabrio
Cluely
Genio
Commure
Retell
CallRail
LiveKit
EliseAI
ClickUp
HeyGen
Models

Nine models, One API.

Enable any combination of features with a single request. Mix, match, and scale as your application grows.

Speaker identification

Label speakers by name using audio context. Supports custom profiles and multi-speaker conversations.

Voice Agents AI Notetaker Call Analytics Agent Assist AI Scribe Conversation Intelligence

Sentiment analysis

Detect positive, neutral, or negative sentiment at the sentence level for granular conversation analytics.

Voice Agents AI Notetaker Call Analytics Agent Assist Conversation Intelligence

Key phrases

Automatically extract the most significant concepts and phrases from every transcript.

Media Monitoring AI Notetaker Call Analytics Agent Assist Conversation Intelligence

Summarization

Automatically extract the most significant concepts and phrases from every transcript.

AI Scribe AI Notetaker Call Analytics Conversation Intelligence

Custom formatting

Normalize dates, phone numbers, and email addresses to machine-readable formats automatically.

Agent Assist AI Notetaker Call Analytics Voice Agents Conversation Intelligence

Translation

Transcribe in any of 99+ languages in the same API request.

Voice Agents AI Notetaker Agent Assist AI Scribe Conversation Intelligence

Auto chapters

Break long recordings into timestamped sections with auto-generated summaries for each chapter.

Content Creation AI Notetaker Conversation Intelligence

Entity detection

Identify names, organizations, locations, dates, and other entities across every transcript.

Voice Agents AI Notetaker Call Analytics Agent Assist AI Scribe Conversation Intelligence

Topic detection

Classify audio content against the IAB standard taxonomy — ideal for content moderation, routing, and analytics.

Media Monitoring AI Notetaker Call Analytics Agent Assist AI Scribe Conversation Intelligence

The second pillar of your Voice AI pipeline

Results 2×

Conversion improvement for teams using structured conversation intelligence vs. raw transcripts alone.

Efficiency 90%

Reduction in manual QA review time after automating complaint and sentiment detection workflows.

Speed <60s

From audio file to fully structured intelligence — speaker labels, sentiment, entities, summary — in one request.

Intelligence 9+

Models for transcription, summarization, sentiment, entities, and more — all in one API.

Speaker labels

Know who said what

  • label

    Label voices by name or role

  • lab_profile

    Attribute summaries to real people

  • person

    Flag compliance issues by speaker

  • hourglass_empty

    Cut hours of manual call review

Sentiment + topics + entities

Surface meaning automatically

  • detector

    Detect sentiment at the sentence level

  • chip_extraction

    Extract named entities automatically

  • quick_phrases

    Pull key phrases that matter most

  • content_copy

    Classify content by IAB topic

Summaries + chapters + formatting

Structure every conversation

  • data_table

    Turn audio into structured data

  • screen_record

    Break recordings into timestamped chapters

  • layers

    Condense calls into one-paragraph briefs

  • format_size

    Format dates, numbers, and contacts

Translation + languages

Go global without switching vendors

  • translate

    Translate 99+ languages in one API

  • audio_video_receiver

    Run sentiment and entities on multilingual audio

  • skip_next

    Skip separate translation vendors

  • analytics

    Run analysis on multilingual audio in one pass

AI Speech-to-Text transcription in 99 languages

From Spanish to Korean, deliver accurate Voice AI in the languages your users speak.

Common questions