New Universal-3.6 Pro Realtime is now available Learn more
Dictation app architecture

The dictation pipeline, end to end

A dictation feature is six steps between a keypress and finished text: capture, bound the clip, pre-warm the connection, upload while the user is still speaking, transcribe and rewrite, then insert. The Dictation API collapses transcription and cleanup into one of them. Here is what each step costs you, and what fails.

Built on the Dictation API rather than on transcription alone:

  • One call returns the transcript and the rewrite
  • A documented five-second bound on the rewrite
  • Constraints you can design against, not discover
  1. 01 Capture on keydown
  2. 02 Bound the clip at 120 s
  3. 03 Pre-warm the connection
  4. 04 Upload while recording
  5. 05 Transcribe + rewrite, one call
  6. 06 Insert the finished text
Metaview
Dovetail
Granola
Apollo.io
Ashby
Siro
Calabrio
Cluely
Genio
Commure
Retell
CallRail
LiveKit
EliseAI
ClickUp
HeyGen
Metaview
Dovetail
Granola
Apollo.io
Ashby
Siro
Calabrio
Cluely
Genio
Commure
Retell
CallRail
LiveKit
EliseAI
ClickUp
HeyGen
Metaview
Dovetail
Granola
Apollo.io
Ashby
Siro
Calabrio
Cluely
Genio
Commure
Retell
CallRail
LiveKit
EliseAI
ClickUp
HeyGen
Metaview
Dovetail
Granola
Apollo.io
Ashby
Siro
Calabrio
Cluely
Genio
Commure
Retell
CallRail
LiveKit
EliseAI
ClickUp
HeyGen

Where the time goes, and what breaks

Step
What happens
Latency effect
Failure mode
1. Capture
Your audio stack records on keydown. WAV or raw 16-bit PCM.
Not on our clock — but start it on keydown, not on key-up.
Microphone permission denied, or a format we reject.
2. Bound the clip
Cap the recording at 120 seconds, the per-call maximum.
A hard ceiling, not a target. Most utterances are seconds long.
Audio over 120 s is rejected. Split it or use a different endpoint.
3. Pre-warm
Call warm() before the user starts speaking.
Takes DNS, TCP and TLS setup off the critical path.
Skip it and the handshake happens inside the user’s wait.
4. Upload while recording
Send the config part first, then audio chunks as you capture them.
Most of the clip is already uploaded when they release the key.
No replay on a chunked body — you cannot retry mid-stream.
5. Transcribe + rewrite
One call returns the verbatim transcript and the rewrite together.
Typically under a second. Every response carries request_time_ms.
Rewrite past five seconds returns HTTP 200 with the verbatim text and llm_error: "timeout".
6. Insert
Write the finished text into whatever has focus.
Nothing left to wait for.
Read final_text, which falls back to the transcript when there is no rewrite.

The constraints, in one place

Short-clip dictation has hard edges. These are the ones worth knowing before you design around them rather than after.

Constraint
Behaviour
Host
dictation.assemblyai.com — its own hostname, separate from pre-recorded, streaming and Sync.
Auth
Your existing API key, sent raw with no Bearer prefix. Different from the rest of the platform.
Audio format
WAV or raw 16-bit PCM only. Compressed audio is rejected with a 415.
Clip length
Up to 120 seconds per call.
Part order
The config part must arrive before the first audio byte.
Retries
No replay on a chunked body.
Rewrite deadline
Five seconds, best-effort. Past it you get the transcript and llm_error, not an error status.
stt_prompt
Up to 6,000 characters.
keyterms_prompt
Up to 100 terms, 8,000 characters total.
llm_instruction
Up to 2,048 characters. Replaces the default cleanup rather than adding to it.
Languages
19 supported languages. The endpoint accepts 32 ISO codes and rejects anything else with a 400 listing the set.
SDK
Python today. Every other language calls the endpoint over plain HTTP as one POST.

Frequently asked questions