Step
What happens
Latency effect
Failure mode
1. Capture
Your audio stack records on keydown. WAV or raw 16-bit PCM.
Not on our clock — but start it on keydown, not on key-up.
Microphone permission denied, or a format we reject.
2. Bound the clip
Cap the recording at 120 seconds, the per-call maximum.
A hard ceiling, not a target. Most utterances are seconds long.
Audio over 120 s is rejected. Split it or use a different endpoint.
3. Pre-warm
Call warm() before the user starts speaking.
Takes DNS, TCP and TLS setup off the critical path.
Skip it and the handshake happens inside the user’s wait.
4. Upload while recording
Send the config part first, then audio chunks as you capture them.
Most of the clip is already uploaded when they release the key.
No replay on a chunked body — you cannot retry mid-stream.
5. Transcribe + rewrite
One call returns the verbatim transcript and the rewrite together.
Typically under a second. Every response carries request_time_ms.
Rewrite past five seconds returns HTTP 200 with the verbatim text and llm_error: "timeout".
6. Insert
Write the finished text into whatever has focus.
Nothing left to wait for.
Read final_text, which falls back to the transcript when there is no rewrite.