Skip to main content
AssemblyAI provides Voice AI infrastructure: pre-recorded speech-to-text, streaming speech-to-text, synchronous speech-to-text, voice agent orchestration, speech understanding, guardrails and LLM access. This breaks down into a few core APIs, suited to different parts of your voice product stack.

Choose how you build

API reference

Build directly against the raw HTTP and WebSocket API for full control over requests and responses.

SDK

Use the Python or TypeScript SDK to integrate transcription, streaming, and voice agents directly into your own codebase.

Build with an AI coding agent

Wire AssemblyAI into Claude Code, Cursor, or any MCP-compatible agent via the AssemblyAI MCP Server and skill.

Integrations

Drop AssemblyAI into existing voice-agent frameworks and platforms — LiveKit, Pipecat, Twilio, Langflow, and others.