Overview
Streaming Diarization lets you identify and label individual speakers in real time directly from the Streaming API. EachTurn event includes a speaker_label
field (e.g. A, B) indicating the dominant speaker for that turn. Each final
word in the words array also carries a speaker field, enabling mid-turn speaker
change detection. Speaker accuracy improves over the course of a session as the model
accumulates embedding context — so the longer the conversation, the better the labels.
Quickstart
Enable Streaming Diarization by settingspeaker_labels to true when you open the WebSocket.
- Python
- Python SDK
- Javascript
- JavaScript SDK
Configuration
Enable Streaming Diarization by addingspeaker_labels: true to your connection
parameters. You can optionally cap the number of speakers with max_speakers.
This is a hard limit, not a hint: if more people speak than the value you set, the
additional speakers are merged into the closest existing speaker label rather than
given a new one. If you’re unsure of the exact speaker count, set max_speakers a
little higher than what you expect so the model has room to identify any additional
speakers. Setting it too high, though, can cause the model to over-split — assigning
new speaker labels to segments that actually belong to an existing speaker.
Diarization is supported on all streaming models:
universal-3-6-pro,
universal-streaming-english, and universal-streaming-multilingual. You do
not need to change your speech model to use it — just add
speaker_labels: true.Reading speaker labels
When diarization is enabled, everyTurn event includes a speaker_label field
reflecting the dominant speaker for that turn, and a speaker_confidence field
scoring how clearly that turn matched the speaker it was assigned to.
Word-level speaker labels
Each final word in thewords array also carries a speaker field. This allows
you to detect speaker changes within a single turn — for example, a turn where one
speaker finishes another’s sentence, or where a brief interjection appears mid-turn.
speaker:
- Final words only. The
speakerfield only appears on words whereword_is_final: true. Non-final (in-progress) words never carry it. speakercan be absent on individual words. If the field is missing from a word entirely, treat that word as unattributed and fall back to the turn-levelspeaker_labelif you need a label. Absent means the field is omitted from the JSON — it will never benull.PENDINGat word level means the model couldn’t confidently attribute that word to any specific speaker — common for short backchannels (“uh huh”, “yeah”) or brief low-quality audio segments. It is not an ambiguity flag between two known speakers; words in a confidently-attributed stretch carry the speaker’s letter, notPENDING.- Each final word also carries a
speaker_confidencescoring how clearly it matched the speaker it was assigned to. See Speaker confidence.
speaker_label
will be set to "PENDING". This is because the model needs at least ~1 second of
audio to generate a reliable diarization embedding — without enough audio, embeddings
may be inaccurate and could lead to a single speaker being labeled as multiple
speakers. Labeling short turns as "PENDING" ensures that speaker labels remain
as accurate as possible.
Speaker confidence
Every finalTurn also carries a speaker_confidence, and so does each of its
words. The score is between 0 and 1 and reflects how clearly the audio matched
the speaker it was assigned to compared with the next-closest speaker: higher
means the two speakers were easy to tell apart, lower means the segment sat
close to the boundary between them. Nothing extra needs to be enabled —
speaker_labels: true is enough.
Keep these details in mind:
- It can be missing. Like
speaker, the field is omitted from the JSON rather than sent asnull. Expect it to be absent on the first words of a session, before the model has built a speaker profile. - Revisions do not carry it.
SpeakerRevisionmessages contain no confidences. Once a revision changes a turn’s labels, any confidence you stored describes the label it replaced — discard it rather than carrying it over to the new one. - Treat a turn-level
0.0as “not available”, not as “certainly wrong”. It appears in rare cases early in a session when no word in the turn contributed a score.
How speaker accuracy improves over time
Streaming Diarization builds a speaker profile incrementally as audio flows in. In practice this means:- Early in a session, speaker assignments may be less stable, especially if the first few turns are short.
- As the session progresses, the model accumulates richer speaker embeddings and assignments become more consistent.
speaker_labels_revision_interval_ms parameter only
controls whether corrections for turns you already received are sent
mid-session.
Revised speaker labels
During a live session, Streaming Diarization assignsspeaker_label values in
real time as each Turn is emitted. These labels can shift as the session
progresses and more audio becomes available. Early turns in particular may
be reassigned as the model builds a clearer picture of each speaker.
When the session ends, the server performs a final refinement pass with full
visibility into the entire conversation. Any turns whose speaker labels changed
are sent back in a SpeakerRevision message. Turns that were already correct
are omitted.
By default that end-of-session message is the only revision you receive. Set
speaker_labels_revision_interval_ms to also receive revisions while the
session is running, so a long transcript can be corrected as it is built rather
than only at the end.
Use the revised labels whenever you need the highest-quality speaker attribution
for the final transcript — for example, when persisting a meeting transcript,
generating a post-call summary, or feeding text into a downstream LLM that
benefits from accurate speaker turns.
The end-of-session refinement adds approximately 400ms of latency. Every
SpeakerRevision message arrives before the Termination message, and none
of them change the real-time speaker_label values already delivered during
the session.Message shape
EverySpeakerRevision message carries a revisions array. Each item corrects
one turn; only turns whose speaker assignments changed since the previous
message are included. The turn_order field in each item matches the
turn_order of the original Turn message being revised. The shape is the same
whether the message arrives mid-session or at the end of the session.
Text content and word timestamps are never changed. Only speaker
assignments are revised.
How to handle it
Match eachturn_order in revisions against the turn you already received,
then replace its speaker_label and per-word speaker values. Keep your turns
in a map keyed by turn_order, as the handlers below do, and the same code
works whether you receive one revision or many.
Each message is a delta: it carries only the turns whose labels changed since
the previous message. Apply items in the order they arrive and let the last
write for a given turn_order win. A turn that you skip is not re-sent in a
later message, so it keeps the label you last stored for it.
- Python
- Python SDK
- JavaScript
When it is sent
A session always ends with a finalSpeakerRevision, sent after the
Terminate signal and before Termination. It includes only the turns whose
speaker assignments changed.
To also receive revisions mid-session, set
speaker_labels_revision_interval_ms to the number of milliseconds of audio
you want between them. We recommend 300000 (5 minutes):
- With a value greater than
0, the server sends a revision roughly every interval of audio, and the first one no sooner than about 2 minutes of audio. The interval counts audio time, not wall-clock time, so silence and under-delivered audio push revisions later. - Values above the server default of
300000are clamped to it, so revisions may arrive more often than you requested. - With
0or the parameter unset, you receive only the end-of-session revision.
Known limitations
Real-time diarization is an inherently harder problem than diarization for async transcription on pre-recorded audio. The following limitations apply to the current beta:- Short utterances — Turns with less than ~1 second of audio are labeled
as
"PENDING"because there is insufficient audio to generate a reliable speaker embedding. This prevents inaccurate embeddings from causing a single speaker to be split across multiple labels. - Overlapping speech — When two speakers talk simultaneously, the model cannot split the audio and will assign the turn to a single speaker. Performance degrades with frequent cross-talk.
- Session start accuracy — The first 1–2 turns of a session may be misassigned because the model has not yet built up speaker profiles. This self-corrects quickly in practice.
- Noisy environments — Background noise and microphone bleed between speakers can reduce embedding quality and lead to more frequent misassignments.
Multichannel streaming audio
To transcribe multichannel streaming audio, we recommend creating a separate session for each channel. This approach allows you to maintain clear speaker separation and get accurate diarized transcriptions for conversations, phone calls, or interviews where speakers are recorded on two different channels. The following code example demonstrates how to transcribe a dual-channel audio file with diarized, speaker-separated transcripts. This same approach can be applied to any multi-channel audio stream, including those with more than two channels.- Python
- Python SDK
- JavaScript
- JavaScript SDK
1
Firstly, install the required dependencies.
2
Use this complete script to transcribe dual-channel audio with speaker separation:
Configure turn detection for your use caseThe examples above use turn detection settings optimized for short responses and rapid back-and-forth conversations. To optimize for your specific audio scenario, you can adjust the turn detection parameters.For configuration examples tailored to different use cases, refer to our Configuration examples.
- Python
- Python SDK
- JavaScript
- JavaScript SDK
Modify the turn detection parameters in
API_PARAMS: