New Universal-3.6 Pro Realtime is now available Learn more
Voice Standard Bench

Choose the right model for every voice task

Shortlist the right LLM for your transcripts faster. Voice Standard Bench tests every LLM Gateway model on real voice tasks, giving you a starting point for benchmarks of your own.

Highest Quality

94

Score, 0-100

GPT-6 Astra
Lowest cost

$0.003

Per hour of audio

Nemotron 3 Nano 30B A3B
Fastest

0.60 s

Median response time

Qwen3.5 4B Fast
The results

Models by output quality

Every Gateway model, scored on real transcripts. Pick a task and transcript length above to see quality, cost per hour of audio and response time.

Quality0 to 100, 2-minute audio duration.
Higher is better
100
GPT-6 Astra
94
GPT-6.1 Sol
93
Opus 5.5
93
Opus 5
93
GPT-5
93
GPT-6 Luna
92
GPT-5.6 Luna
92
DeepSeek V4.1 Flash
92

Overall quality score by model on 2-minute transcripts, highest first. Showing the top 8 of 50+ models available in LLM Gateway.

How the quality score is built ↗

Price performance
Cost per hour of audio, USD
Best price performance1009590858075706560$0.005$0.010$0.050$0.100$0.500Nemotron 3 Nano 30B A3BGemini 2.5 Flash Litegemma-4-31bGPT-6 LunaGPT-6.1 SolOpus 5.5Opus 5GPT-6 AstraGPT-5

Quality score against cost per hour of audio overall on 2-minute transcripts. The band marks the models with the best price performance. Showing 9 of 50+ models.

How the quality score is built ↗

Cost per hour of audio
Cost per hour of audio, USD
  • GPT-6 Luna$0.008
  • GPT-5.6 Luna$0.015
  • DeepSeek V4.1 Flash$0.073
  • GPT-6.1 Sol$0.119
  • Opus 5.5$0.394
  • Opus 5$0.439
  • GPT-6 Astra$0.505
  • GPT-5$0.666

Cost per hour of audio overall on 2-minute transcripts, cheapest first, among the top 8 models by quality.

Response time
Average response time, seconds
  • GPT-5.6 Luna3.20 s
  • GPT-6 Luna3.41 s
  • GPT-6 Astra5.42 s
  • Opus 55.52 s
  • GPT-6.1 Sol6.42 s
  • Opus 5.59.89 s
  • GPT-522.53 s
  • DeepSeek V4.1 Flash32.25 s

Average response time per call on 2-minute transcripts, fastest first, among the top 8 models by quality.

Head to head

Compare the models you are already considering

Pick three Gateway models and see quality, cost per hour of audio and median response time side by side.

Model A
Model B
Model C
DimensionOpus 5Qwen3.5 4B FastGPT-5.6 Luna
Quality936492
Cost per hour of audio$0.439$0.005$0.015
Median response time5.02 seconds0.60 seconds2.20 seconds
QualityQuality, 0 to 100
  • GPT-6 Astra94
  • GPT-6.1 Sol93
  • Opus 5.593
  • Opus 593
  • GPT-593
  • GPT-6 Luna92
  • GPT-5.6 Luna92
  • Qwen3.5 4B Fast64
Cost per hour of audioCost per hour of audio, USD
  • Qwen3.5 4B Fast$0.005
  • GPT-6 Luna$0.008
  • GPT-5.6 Luna$0.015
  • GPT-6.1 Sol$0.119
  • Opus 5.5$0.394
  • Opus 5$0.439
  • GPT-6 Astra$0.505
  • GPT-5$0.666
Median response timeMedian response time, seconds
  • Qwen3.5 4B Fast0.60 s
  • GPT-5.6 Luna2.20 s
  • GPT-6 Luna2.37 s
  • GPT-6 Astra3.76 s
  • Opus 55.02 s
  • Opus 5.55.38 s
  • GPT-6.1 Sol6.76 s
  • GPT-517.81 s
Run your own

Benchmark with your own data in LLM Gateway

These scores are a starting point. Run the models that look right on your own transcripts before you decide.

Get the JSON

Every figure on this page, for every model, task and transcript length, as JSON on GitHub.

Open on GitHub

Test on your own transcripts

Create a free account and run any model here on your transcripts, with the same API key as your Speech-to-Text.

Start building

Common questions