# Real-time Speech-to-Text API

AssemblyAI's Real-time Speech-to-Text API transcribes live audio streams with low latency and production accuracy.

Canonical URL: https://www.assemblyai.com/products/streaming-speech-to-text

## Summary

Use streaming transcription for voice agents, live captions, agent assist, meetings, phone calls, and any workflow where audio is processed while it is happening. Billing is based on WebSocket session duration.

## Core Models

- Universal-3.5 Pro Realtime: highest-accuracy real-time model for voice agents and critical live workflows.
- Universal-Streaming: lower-cost real-time transcription for English.
- Universal-Streaming Multilingual: lower-cost real-time transcription across supported multilingual workflows.
- Whisper-Streaming: real-time option for broader language coverage.

## Common Use Cases

- Voice agents and conversational AI.
- Live captions and accessibility.
- Agent assist for contact centers.
- Real-time meeting transcription.
- Live analytics over calls and events.

## Related References

- https://www.assemblyai.com/docs/streaming/getting-started/transcribe-streaming-audio.md
- https://www.assemblyai.com/docs/streaming/select-the-speech-model.md
- https://www.assemblyai.com/docs/streaming/universal-3-pro.md
- https://www.assemblyai.com/docs/streaming/keyterms-prompting.md
- https://www.assemblyai.com/docs/streaming/label-speakers-and-separate-channels.md
- https://www.assemblyai.com/pricing.md
