New Universal-3.5 Pro is here. Learn more: Async Realtime
Features

Every feature, on one API request

Build products that know who spoke, what mattered, and what to keep private. Enable any combination of features on the same transcription request—no extra pipelines to build, no models to maintain.

Universal-3.5 Pro

Upload a file to see transcription and Speech Understanding in action.

Metaview
Dovetail
Granola
Apollo.io
Ashby
Siro
Calabrio
Cluely
Genio
Commure
Retell
CallRail
LiveKit
Earmark
ClickUp
HeyGen
Metaview
Dovetail
Granola
Apollo.io
Ashby
Siro
Calabrio
Cluely
Genio
Commure
Retell
CallRail
LiveKit
Earmark
ClickUp
HeyGen

Get production-ready outputs in a single API call

Every feature below runs on the same transcription request. Pick the ones your product needs and add them with a single parameter.

Speaker Diarization

Separate every voice in a conversation into a structured, speaker-labeled transcript with timestamps.

Speaker Identification

Replace “Speaker A” and “Speaker B” with real names or roles, inferred from the conversation itself.

Summarization + Chapters

Turn every recording into concise summaries and timestamped chapters, each with a headline and gist.

Action Items

Pull the follow-ups out of any conversation, each with a timestamp and the quote it came from.

Sentiment Analysis

Classify sentiment as positive, neutral, or negative, sentence by sentence, and per speaker when diarization is on.

Key Phrases

Surface the words and phrases that matter most in a recording, ranked by relevance with timestamps.

Entity Detection

Get names, dates, account numbers, and 50+ other entity types back as typed fields with timestamps.

Topic Detection

Classify what audio is about against the standard IAB Content Taxonomy categories, at scale.

Custom Formatting

Standardize dates, phone numbers, currency, and URLs in the transcript itself, ready for the systems downstream.

Automatic Language Detection

Detect the dominant language in any audio file and transcribe it across 99 languages, from one API.

Translation

Translate transcripts into 86 languages in the same request, in a formal or informal style.

PII Redaction

Remove personally identifiable information from the transcript, and bleep it from the audio itself, in a single API call.

Content Moderation

Detect sensitive content and clean up offensive language automatically, before it ships.

Built to scale

Enterprise-grade features, infrastructure built for scale

Every feature runs on the same industry-leading Voice AI models and inference platform behind 800M+ API calls a month, with the accuracy, availability, and compliance controls enterprise deployments require.

Every feature runs on our Voice AI models

4.35% word error rate on Universal-3.5 Pro

Sub-300ms latency for streaming and voice agents

800M+ API calls processed every month

Coverage across 99 languages

No concurrency limits or throttles at any scale

SOC 2 Type 2 and GDPR, with BAA available

Deploy in our cloud or self-hosted in your environment

Common questions