New Universal-3.5 Pro is here. Learn more: Async Realtime
LLM Gateway

Run open and frontier LLMs on your voice data

From raw utterance to structured answer—fast open models for the cleanup pass, frontier models for the reasoning pass, one API key for both.

Speech-to-Text

Hi, Marissa. Um, quick recap on where we landed, uh, with the Q3 vendor review. So, uh, we got through 6 of the 8— oh no, 5 of the 8 assessments. Um, and the last 3, they’re waiting on security sign-off. Um, the big thing is that pricing came back higher than, than we modeled, something like, uh, between 12— no, 15% over. So I think we need to revisit The budget line before we commit to anything. Can you pull the original forecast for me? And I’ll put 30 minutes on the calendar.

LLM Gateway

Clean up dictated speech. Remove fillers (um, uh, like, you know), stutters, repetitions, false starts; on self-corrections keep only the final version. Fix punctuation, capitalization, sentence and paragraph breaks. Render spoken commands ("period", "new paragraph") as punctuation.

Final output

Hi Marissa, quick recap on where we landed with the Q3 vendor review. We got through five of the eight assessments. The last three are waiting on security sign-off. The big thing is that pricing came back higher than we modeled, something like 15% over. I think we need to revisit the budget line before we commit to anything. Can you pull the original forecast for me? I'll put 30 minutes on the calendar.

539 ms · 90 tokens · <$0.0001
Metaview
Dovetail
Granola
Apollo.io
Ashby
Siro
Calabrio
Cluely
Genio
Commure
Retell
CallRail
LiveKit
Earmark
ClickUp
HeyGen
Metaview
Dovetail
Granola
Apollo.io
Ashby
Siro
Calabrio
Cluely
Genio
Commure
Retell
CallRail
LiveKit
Earmark
ClickUp
HeyGen
Models

Open and frontier models, one endpoint

34 models from 5 providers, all OpenAI-compatible—switch by changing one string.

Provider
Input rate
LLM Gateway models with provider and input and output token rates per 1 million tokens.
Models Provider Input Output
GPT-5 Nano Open AI $0.05 / 1M $0.40 / 1M
GPT OSS 20B Bedrock $0.07 / 1M $0.30 / 1M
Gemini 2.5 Flash Lite Vertex $0.10 / 1M $0.40 / 1M
Qwen3.5 4B Fast AssemblyAI $0.10 / 1M $0.50 / 1M
gemma-4-31b Bedrock Mantle $0.14 / 1M $0.40 / 1M
GPT OSS 120B Bedrock $0.15 / 1M $0.60 / 1M
Qwen3 32B Bedrock $0.15 / 1M $0.60 / 1M
Qwen3 Next 80B A3B Bedrock $0.15 / 1M $1.20 / 1M
Gemini 3.1 Flash Lite Vertex $0.25 / 1M $1.50 / 1M
GPT-5 mini Open AI $0.25 / 1M $2.00 / 1M
Gemini 2.5 Flash Vertex $0.30 / 1M $2.50 / 1M
Gemini 3.5 Flash Lite Vertex $0.30 / 1M $2.50 / 1M
Gemini 3.7 Flash Vertex $0.75 / 1M $3.75 / 1M
Gemini 3.8 Flash Vertex $0.75 / 1M $3.75 / 1M
GPT-5.6 Luna Open AI $1.00 / 1M $6.00 / 1M
Haiku 4.5 Bedrock $1.00 / 1M $5.00 / 1M
Gemini 2.5 Pro Vertex $1.25 / 1M $10.00 / 1M
Gemini 3.5 Flash Vertex $1.25 / 1M $9.00 / 1M
GPT-5 Open AI $1.25 / 1M $10.00 / 1M
GPT-5.1 Open AI $1.25 / 1M $10.00 / 1M
Gemini 3.6 Flash Vertex $1.50 / 1M $7.50 / 1M
GPT-5.2 Open AI $1.75 / 1M $14.00 / 1M
GPT-4.1 Open AI $2.00 / 1M $8.00 / 1M
GPT-5.6 Terra Open AI $2.50 / 1M $15.00 / 1M
Sonnet 4.5 Bedrock $3.00 / 1M $15.00 / 1M
Sonnet 4.6 Bedrock $3.00 / 1M $15.00 / 1M
Sonnet 5 Bedrock $3.00 / 1M $15.00 / 1M
GPT-5.6 Sol Open AI $4.00 / 1M $20.00 / 1M
GPT-5.5 Open AI $5.00 / 1M $30.00 / 1M
Opus 4.5 Bedrock $5.00 / 1M $25.00 / 1M
Opus 4.6 Bedrock $5.00 / 1M $25.00 / 1M
Opus 4.7 Bedrock $5.00 / 1M $25.00 / 1M
Opus 4.8 Bedrock $5.00 / 1M $25.00 / 1M
Opus 5 Bedrock $5.00 / 1M $25.00 / 1M

Showing 34 of 34 models

Prices shown are for global routing. In-region (US/EU) pricing is 10% higher due to provider cost increases. Every model, with context windows and capabilities, is on the full catalog.

Production grade

Easiest, most reliable way to call multiple LLMs

Ship faster, spend less on tokens, and stop losing users to provider outages.

0% markup

Pay provider rates, not gateway rates. Competing gateways add 5% or more to every call.

Automatic fallbacks

Configure backup models per request. When a provider errors or stalls, your call still goes through.

Security by simplicity

Pay the exact same price as calling the model provider directly. No markup, no hidden fees, no minimum commitment. We make it simple.

OpenAI-compatible

Drop into any OpenAI SDK. Change a base URL and a model string, and everything keeps working.

Voice-native

Your LLM calls run where your transcription does. One less network hop on every turn.

Models worth using

Frontier models from OpenAI, Anthropic, and Google, plus open models we host ourselves. New ones added the day they launch.

Common questions