New Universal-3.5 Pro is here. Learn more: Async Realtime
LLM Gateway

Browse LLM Gateway models

Compare context, pricing, and capabilities across every model the Gateway serves, and run any of them with the same API key as your Speech-to-Text.

Providers0
Regions0
Features0
New models first

Showing 34 of 34 models

LLM Gateway models with context window, input, output and cached token rates per 1 million tokens, capabilities, serving providers and regions.
Capabilities Providers Regions
anthropic/claude-opus-5 New
200K $5.00/M $25.00/M
Read:$0.50/M Write:$6.25/M
Tools Streaming Bedrock USGlobal
google/gemini-3.8-flash New
1M $0.75/M $3.75/M
Read:$0.07/M Write:
JSON Tools Streaming Vertex USEUGlobal
openai/gpt-5.6-sol New
270K $4.00/M $20.00/M
Read:$0.40/M Write:$5.00/M
JSON Tools Streaming Open AI USGlobal
google/gemma-4-31b New
256K $0.14/M $0.40/M
Read: Write:
JSON Tools Streaming Bedrock Mantle US
google/gemini-3.7-flash New
1M $0.75/M $3.75/M
Read:$0.07/M Write:
JSON Tools Streaming Vertex USEUGlobal
openai/gpt-5-nano
400K $0.05/M $0.40/M
Read:$0.005/M Write:
JSON Tools Streaming Open AI US
openai/gpt-oss-20b
131K $0.07/M $0.30/M
Read: Write:
Tools Bedrock US
alibaba/qwen3.5-4b-32k-fast
33K $0.10/M $0.50/M
Read: Write:
Streaming AssemblyAI USEU
google/gemini-2.5-flash-lite
1M $0.10/M $0.40/M
Read:$0.01/M Write:
JSON Tools Streaming Vertex USEUGlobal
alibaba/qwen3-32B
200K $0.15/M $0.60/M
Read: Write:
Tools JSON Streaming Bedrock US
alibaba/qwen3-next-80b-a3b
200K $0.15/M $1.20/M
Read: Write:
Tools JSON Streaming Bedrock US
openai/gpt-oss-120b
131K $0.15/M $0.60/M
Read: Write:
JSON Tools Bedrock US
google/gemini-3.1-flash-lite
1M $0.25/M $1.50/M
Read:$0.03/M Write:
JSON Tools Streaming Vertex USGlobal
openai/gpt-5-mini
400K $0.25/M $2.00/M
Read:$0.03/M Write:
JSON Tools Streaming Open AI US
google/gemini-2.5-flash
1M $0.30/M $2.50/M
Read:$0.03/M Write:
JSON Tools Streaming Vertex USEUGlobal
google/gemini-3.5-flash-lite
1M $0.30/M $2.50/M
Read:$0.03/M Write:
JSON Tools Streaming Vertex USGlobal
anthropic/claude-haiku-4-5-20251001
200K $1.00/M $5.00/M
Read:$0.10/M Write:$1.25/M
Tools JSON Streaming Bedrock USEUGlobal
openai/gpt-5.6-luna
270K $1.00/M $6.00/M
Read:$0.10/M Write:$1.25/M
JSON Tools Streaming Open AI USGlobal
google/gemini-2.5-pro
200K $1.25/M $10.00/M
Read:$0.13/M Write:
JSON Tools Streaming Vertex USEUGlobal
google/gemini-3.5-flash
1M $1.25/M $9.00/M
Read:$0.13/M Write:
JSON Tools Streaming Vertex USGlobal
openai/gpt-5
400K $1.25/M $10.00/M
Read:$0.13/M Write:
Tools Streaming JSON Open AI US
openai/gpt-5.1
400K $1.25/M $10.00/M
Read:$0.13/M Write:
JSON Tools Streaming Open AI US
google/gemini-3.6-flash
1M $1.50/M $7.50/M
Read:$0.15/M Write:
JSON Tools Streaming Vertex USEUGlobal
openai/gpt-5.2
400K $1.75/M $14.00/M
Read:$0.17/M Write:
JSON Tools Streaming Open AI US
openai/gpt-4.1
1M $2.00/M $8.00/M
Read:$0.50/M Write:
Tools Streaming Open AI US
openai/gpt-5.6-terra
270K $2.50/M $15.00/M
Read:$0.25/M Write:$3.13/M
JSON Tools Streaming Open AI USGlobal
anthropic/claude-sonnet-4-5-20250929
200K $3.00/M $15.00/M
Read:$0.30/M Write:$3.75/M
Tools JSON Streaming Bedrock USEUGlobal
anthropic/claude-sonnet-4-6
200K $3.00/M $15.00/M
Read:$0.30/M Write:$3.75/M
Tools JSON Streaming Bedrock USEUGlobal
anthropic/claude-sonnet-5
200K $3.00/M $15.00/M
Read:$0.30/M Write:$3.75/M
Tools Streaming Bedrock USGlobal
anthropic/claude-opus-4-5-20251101
200K $5.00/M $25.00/M
Read:$0.50/M Write:$6.25/M
Tools JSON Streaming Bedrock USGlobal
anthropic/claude-opus-4-6
200K $5.00/M $25.00/M
Read:$0.50/M Write:$6.25/M
Tools JSON Streaming Bedrock USGlobal
anthropic/claude-opus-4-7
1M $5.00/M $25.00/M
Read:$0.50/M Write:$6.25/M
Tools Streaming Bedrock USGlobal
anthropic/claude-opus-4-8
1M $5.00/M $25.00/M
Read:$0.50/M Write:$6.25/M
Tools Streaming Bedrock US
openai/gpt-5.5
272K $5.00/M $30.00/M
Read:$0.50/M Write:
JSON Tools Streaming Open AI USGlobal

Models are listed as maker/model. The part after the slash is the exact model value to pass in a request. The Gateway does not accept the maker prefix.

Rates are for global routing. In-region US and EU endpoints are 10% higher, passed through from provider pricing with no AssemblyAI upcharge. See pricing for the full rate card.

Cache read is the rate for tokens served from the prompt cache; cache write is the rate to store them. An em dash means the model has no cache tier. See prompt caching.

Providers are where a request is served from — Claude models run on Bedrock, Gemini on Vertex, and Qwen3.5 4B Fast on AssemblyAI's own GPUs. This table is generated from the Gateway's own Models endpoint, so it always matches what the API accepts. The docs roster adds quality scores and measured latency.