New Universal-3.6 Pro Realtime is now available Learn more
Dictation for coding agents

Dictate the prompt, not the code

Coding agents take their instructions in prose, which makes them a natural fit for voice — but a verbatim transcript is a poor prompt. The Dictation API returns the instruction you meant: self-corrections resolved, filler gone, your repo and library names spelled right, in a single call.

One API call between the microphone and the agent:

  • Transcription plus an LLM cleanup pass, together
  • A rewrite instruction you control per request
  • Finished text back in under a second

What you said

okay so um in the the auth module — no wait, the middleware — add a rate limiter, like sixty requests a minute, and uh make sure it uses redis not in-memory because we’re running like four instances

What the agent receives

In the middleware, add a rate limiter of 60 requests per minute. Use Redis rather than in-memory storage, since we run four instances.

Metaview
Dovetail
Granola
Apollo.io
Ashby
Siro
Calabrio
Cluely
Genio
Commure
Retell
CallRail
LiveKit
EliseAI
ClickUp
HeyGen
Metaview
Dovetail
Granola
Apollo.io
Ashby
Siro
Calabrio
Cluely
Genio
Commure
Retell
CallRail
LiveKit
EliseAI
ClickUp
HeyGen
Metaview
Dovetail
Granola
Apollo.io
Ashby
Siro
Calabrio
Cluely
Genio
Commure
Retell
CallRail
LiveKit
EliseAI
ClickUp
HeyGen
Metaview
Dovetail
Granola
Apollo.io
Ashby
Siro
Calabrio
Cluely
Genio
Commure
Retell
CallRail
LiveKit
EliseAI
ClickUp
HeyGen

The instruction

You write the rewrite rule once

The cleanup pass is not a fixed behaviour you have to accept. Send an llm_instruction and the output arrives in the shape you asked for — this is the one that produced the prompt above.

Rewrite dictated developer instructions as a single
imperative prompt for a coding agent. Resolve self-corrections
to the speaker's final choice. Convert spoken numbers to digits.
Keep file, module, and library names exactly as spoken.
Return the prompt only, with no preamble.

It replaces the default cleanup rather than adding to it, so state everything you need. Limits and the full config are in the Dictation API docs.

What the cleanup pass is actually worth

For an agent prompt
Dictation API Transcribe + rewrite, one call
Transcription alone Verbatim transcript
What the agent receives
A finished instruction
Every false start, filler word, and restart
Self-corrections
Resolved to what you landed on
Left in the text for the agent to guess at
Repo, service, and library names
Keyterms prompting, included
Spelled as they sounded
Output shape
Set per request with an instruction
—
Spoken numbers
Digits, per your instruction
"sixty requests a minute"
Calls per utterance
One
Two — transcription, then your own LLM prompt
If the rewrite runs long
Verbatim text returns at the five-second bound, HTTP 200
—
Languages
19, with mid-sentence code-switching
—

Three knobs, all included in the rate

stt_prompt

Describe what the model is about to hear — your stack, your repo, the kind of speech your users produce. Up to 6,000 characters.

keyterms_prompt

Bias the transcript toward the names it would otherwise guess at: your services, libraries, and teammates. Up to 100 terms.

llm_instruction

Describe the shape you want back. It replaces the default cleanup rather than adding to it, so state everything you need. Up to 2,048 characters.

Frequently asked questions