New Universal-3.5 Pro is here. Learn more: Async Realtime
Features

Entity Detection

Turn speech into structured data. Names, dates, account numbers, and medications come back as typed fields with timestamps — not regex you maintain on top of a transcript.

Get started with less than 10 lines of code

Enable Entity Detection in your transcription request and receive a structured list of every entity found in the audio, each with its type and exact position.

Detect 50+ entity types

People, organizations, locations, dates and times, phone numbers, medical conditions, financial details, and more — classified automatically.

Typed fields, not string matching

Every entity is returned with its category and millisecond-accurate timestamps, so you can act on it in code instead of parsing raw text.

Works across 50+ languages

Extract entities from multilingual audio in the same request, powered by our Universal speech-to-text models.

Use cases

Structure every conversation automatically

Pull the who, what, and when out of audio without building your own extraction layer.

Populate CRM fields from calls

Route support tickets by entity

Index media by people and places

Extract order and account numbers

Surface medications in clinical audio

Enrich analytics with entities

Auto-tag meeting notes

Feed downstream LLM workflows

Join 200K+ developers building new experiences with voice data

The model is very promising. Transcript quality and accuracy are noticeably better, especially when it comes to correctly handling participant names and key terms.

GD

Galya Dimitrova

Head of Product, Jiminny

If you have an hour of content, the difference between 99% accuracy and 97% accuracy, it's a lot of time for that person to review. So you could cut down their workflow from taking half an hour, to 20 minutes, to 15 minutes — it's huge, right?

Joshua Grossberg

Joshua Grossberg

CTO, Kapwing

Investments in STT improvements always pay for themselves, since it is such a critical building block of the voice pipeline.

Lindsay Liu

Lindsay Liu

Co-Founder & CEO, Super

AssemblyAI's managed API endpoint and diarization won me over — something Whisper couldn't provide.

Josh Mohrer

Josh Mohrer

Founder, Wave.co

We needed a provider that could scale with us — offering unlimited concurrent streams, fair pricing, and responsive support.

Mark Barbir

Mark Barbir

CEO, Earmark

The transcription accuracy, reliability, and speed of AssemblyAI's API have greatly enhanced our operations.

Raj Shankar

Raj Shankar

SVP Product, Calabrio

Our free to paid conversion rate doubled after implementing AssemblyAI.

Colin Treseler

Colin Treseler

Founder & CEO, Supernormal

The accuracy was strong, but the great documentation and unique models like Auto Chapters and Sentiment Analysis is what really won us over.

Nathan Webb

Nathan Webb

Product Manager, Aloware

Calls are and will remain a pertinent part of the customer service journey. Customer service isn't moving entirely to chatbots or chat interactions.

Dr. Shane Lynn

Dr. Shane Lynn

CEO, EdgeTier