Skip to main content
Identify and label individual speakers in real time, or transcribe multichannel audio using the Streaming API.

Overview

Streaming Diarization lets you identify and label individual speakers in real time directly from the Streaming API. Each Turn event includes a speaker_label field (e.g. A, B) indicating the dominant speaker for that turn. Each final word in the words array also carries a speaker field, enabling mid-turn speaker change detection. Speaker accuracy improves over the course of a session as the model accumulates embedding context — so the longer the conversation, the better the labels.
Already using AssemblyAI streaming?You can enable Streaming Diarization by adding speaker_labels: true to your connection parameters. No other changes are required — the speaker_label field will appear on every Turn event, and each final word in the words array will include a speaker field automatically.

Quickstart

Enable Streaming Diarization by setting speaker_labels to true when you open the WebSocket.

Configuration

Enable Streaming Diarization by adding speaker_labels: true to your connection parameters. You can optionally cap the number of speakers with max_speakers. This is a hard limit, not a hint: if more people speak than the value you set, the additional speakers are merged into the closest existing speaker label rather than given a new one. If you’re unsure of the exact speaker count, set max_speakers a little higher than what you expect so the model has room to identify any additional speakers. Setting it too high, though, can cause the model to over-split — assigning new speaker labels to segments that actually belong to an existing speaker.
Diarization is supported on all streaming models: universal-3-6-pro, universal-streaming-english, and universal-streaming-multilingual. You do not need to change your speech model to use it — just add speaker_labels: true.

Reading speaker labels

When diarization is enabled, every Turn event includes a speaker_label field reflecting the dominant speaker for that turn, and a speaker_confidence field scoring how clearly that turn matched the speaker it was assigned to.

Word-level speaker labels

Each final word in the words array also carries a speaker field. This allows you to detect speaker changes within a single turn — for example, a turn where one speaker finishes another’s sentence, or where a brief interjection appears mid-turn.
A few things to keep in mind when consuming speaker:
  • Final words only. The speaker field only appears on words where word_is_final: true. Non-final (in-progress) words never carry it.
  • speaker can be absent on individual words. If the field is missing from a word entirely, treat that word as unattributed and fall back to the turn-level speaker_label if you need a label. Absent means the field is omitted from the JSON — it will never be null.
  • PENDING at word level means the model couldn’t confidently attribute that word to any specific speaker — common for short backchannels (“uh huh”, “yeah”) or brief low-quality audio segments. It is not an ambiguity flag between two known speakers; words in a confidently-attributed stretch carry the speaker’s letter, not PENDING.
  • Each final word also carries a speaker_confidence scoring how clearly it matched the speaker it was assigned to. See Speaker confidence.
If a turn contains less than approximately 1 second of audio, the turn-level speaker_label will be set to "PENDING". This is because the model needs at least ~1 second of audio to generate a reliable diarization embedding — without enough audio, embeddings may be inaccurate and could lead to a single speaker being labeled as multiple speakers. Labeling short turns as "PENDING" ensures that speaker labels remain as accurate as possible.
Your application should handle this case gracefully. A typical multi-speaker exchange looks like this:

Speaker confidence

Every final Turn also carries a speaker_confidence, and so does each of its words. The score is between 0 and 1 and reflects how clearly the audio matched the speaker it was assigned to compared with the next-closest speaker: higher means the two speakers were easy to tell apart, lower means the segment sat close to the boundary between them. Nothing extra needs to be enabled — speaker_labels: true is enough.
Experimental fieldspeaker_confidence is not calibrated, so it is not the probability that the label is correct. Do not apply absolute thresholds to it. Use it for relative comparison instead — for example, to surface the least-confident regions of a transcript for review. We are improving calibration and welcome your feedback on how the scores behave on your audio.
Keep these details in mind:
  • It can be missing. Like speaker, the field is omitted from the JSON rather than sent as null. Expect it to be absent on the first words of a session, before the model has built a speaker profile.
  • Revisions do not carry it. SpeakerRevision messages contain no confidences. Once a revision changes a turn’s labels, any confidence you stored describes the label it replaced — discard it rather than carrying it over to the new one.
  • Treat a turn-level 0.0 as “not available”, not as “certainly wrong”. It appears in rare cases early in a session when no word in the turn contributed a score.

How speaker accuracy improves over time

Streaming Diarization builds a speaker profile incrementally as audio flows in. In practice this means:
  • Early in a session, speaker assignments may be less stable, especially if the first few turns are short.
  • As the session progresses, the model accumulates richer speaker embeddings and assignments become more consistent.
For long-form use cases (call center, clinical scribe, meeting transcription), the model will settle into accurate, stable labels well before the end of the conversation. The server refines its speaker model periodically on its own, so labels on new turns can keep improving during a session without any extra configuration — the speaker_labels_revision_interval_ms parameter only controls whether corrections for turns you already received are sent mid-session.

Revised speaker labels

During a live session, Streaming Diarization assigns speaker_label values in real time as each Turn is emitted. These labels can shift as the session progresses and more audio becomes available. Early turns in particular may be reassigned as the model builds a clearer picture of each speaker. When the session ends, the server performs a final refinement pass with full visibility into the entire conversation. Any turns whose speaker labels changed are sent back in a SpeakerRevision message. Turns that were already correct are omitted. By default that end-of-session message is the only revision you receive. Set speaker_labels_revision_interval_ms to also receive revisions while the session is running, so a long transcript can be corrected as it is built rather than only at the end. Use the revised labels whenever you need the highest-quality speaker attribution for the final transcript — for example, when persisting a meeting transcript, generating a post-call summary, or feeding text into a downstream LLM that benefits from accurate speaker turns.
The end-of-session refinement adds approximately 400ms of latency. Every SpeakerRevision message arrives before the Termination message, and none of them change the real-time speaker_label values already delivered during the session.

Message shape

Every SpeakerRevision message carries a revisions array. Each item corrects one turn; only turns whose speaker assignments changed since the previous message are included. The turn_order field in each item matches the turn_order of the original Turn message being revised. The shape is the same whether the message arrives mid-session or at the end of the session.
Text content and word timestamps are never changed. Only speaker assignments are revised.

How to handle it

Match each turn_order in revisions against the turn you already received, then replace its speaker_label and per-word speaker values. Keep your turns in a map keyed by turn_order, as the handlers below do, and the same code works whether you receive one revision or many. Each message is a delta: it carries only the turns whose labels changed since the previous message. Apply items in the order they arrive and let the last write for a given turn_order win. A turn that you skip is not re-sent in a later message, so it keeps the label you last stored for it.

When it is sent

A session always ends with a final SpeakerRevision, sent after the Terminate signal and before Termination. It includes only the turns whose speaker assignments changed. To also receive revisions mid-session, set speaker_labels_revision_interval_ms to the number of milliseconds of audio you want between them. We recommend 300000 (5 minutes):
  • With a value greater than 0, the server sends a revision roughly every interval of audio, and the first one no sooner than about 2 minutes of audio. The interval counts audio time, not wall-clock time, so silence and under-delivered audio push revisions later.
  • Values above the server default of 300000 are clamped to it, so revisions may arrive more often than you requested.
  • With 0 or the parameter unset, you receive only the end-of-session revision.
If the session ends unexpectedly (network drop, error closure), the final revision may not be delivered. Always handle this gracefully and fall back to the labels you received during the session.

Known limitations

Real-time diarization is an inherently harder problem than diarization for async transcription on pre-recorded audio. The following limitations apply to the current beta:
  • Short utterances — Turns with less than ~1 second of audio are labeled as "PENDING" because there is insufficient audio to generate a reliable speaker embedding. This prevents inaccurate embeddings from causing a single speaker to be split across multiple labels.
  • Overlapping speech — When two speakers talk simultaneously, the model cannot split the audio and will assign the turn to a single speaker. Performance degrades with frequent cross-talk.
  • Session start accuracy — The first 1–2 turns of a session may be misassigned because the model has not yet built up speaker profiles. This self-corrects quickly in practice.
  • Noisy environments — Background noise and microphone bleed between speakers can reduce embedding quality and lead to more frequent misassignments.
For the best results, use a microphone setup that minimizes cross-talk and background noise, and ensure each speaker produces at least a few complete sentences before you rely on per-turn labels for downstream processing.

Multichannel streaming audio

To transcribe multichannel streaming audio, we recommend creating a separate session for each channel. This approach allows you to maintain clear speaker separation and get accurate diarized transcriptions for conversations, phone calls, or interviews where speakers are recorded on two different channels. The following code example demonstrates how to transcribe a dual-channel audio file with diarized, speaker-separated transcripts. This same approach can be applied to any multi-channel audio stream, including those with more than two channels.
1
Firstly, install the required dependencies.
2
Use this complete script to transcribe dual-channel audio with speaker separation:
Configure turn detection for your use caseThe examples above use turn detection settings optimized for short responses and rapid back-and-forth conversations. To optimize for your specific audio scenario, you can adjust the turn detection parameters.For configuration examples tailored to different use cases, refer to our Configuration examples.
Modify the turn detection parameters in API_PARAMS: