For an agent prompt
Dictation API Transcribe + rewrite, one call
Transcription alone Verbatim transcript
What the agent receives
A finished instruction
Every false start, filler word, and restart
Self-corrections
Resolved to what you landed on
Left in the text for the agent to guess at
Repo, service, and library names
check_circle Keyterms prompting, included
Spelled as they sounded
Output shape
Set per request with an instruction
—
Spoken numbers
Digits, per your instruction
"sixty requests a minute"
Calls per utterance
One
Two — transcription, then your own LLM prompt
If the rewrite runs long
Verbatim text returns at the five-second bound, HTTP 200
—
Languages
19, with mid-sentence code-switching
—