New Universal-3.5 Pro is here. Learn more: Async Realtime
JavaScript speech-to-text API

The speech-to-text API for JavaScript developers

Build transcription into your app with a couple of API calls—transcribe audio files from a Node backend or stream live audio over a WebSocket, and get accurate text back with timestamps and speaker labels. Start free, no credit card.

Plain JavaScript, no lock-in

Call the REST API with fetch from any runtime—Node, Deno, Bun, or the browser. A typed npm package is there if you want it, never required.

Real-time from any stream

Stream mic, telephony, or in-app audio over a WebSocket and partial and final transcripts come back within a few hundred milliseconds.

Audio files to text

Upload a file or point the API at a URL, then poll for the finished transcript—formatted text with timestamps in structured JSON.

Metaview
Ashby
Cluely
Genio
Siro

36%

improvement in close rate

See case study
LiveKit
Earmark
Commure
Dovetail
Fireflies

“The new Universal-3.5 Pro speech model from AssemblyAI is best so far in terms of accuracy, latency, and language switching.”

Retell
CallRail
Apollo.io
ClickUp
Calabrio

80%

increase in customer satisfaction

HeyGen
Granola
Siro
JotPsych
Granola

“Assembly has saved us countless hours managing models, and provided exceptional accuracy.”

Metaview
Ashby
Cluely
Genio
Siro

36%

improvement in close rate

See case study
LiveKit
Earmark
Commure
Dovetail
Fireflies

“The new Universal-3.5 Pro speech model from AssemblyAI is best so far in terms of accuracy, latency, and language switching.”

Retell
CallRail
Apollo.io
ClickUp
Calabrio

80%

increase in customer satisfaction

HeyGen
Granola
Siro
JotPsych
Granola

“Assembly has saved us countless hours managing models, and provided exceptional accuracy.”

Built for applications

A JavaScript API for speech-to-text in your app

Transcribe uploads in a Node batch job, power live captions in a meeting tool, or feed a voice agent from a telephony stream—the same API handles files and real-time audio. Every transcript comes back as structured JSON with speaker labels, timestamps, and language detection, so it drops straight into your product instead of a dictation demo.

Start building free See the live demo in the Playground

No credit card required

Streaming word error rate

Word error rate is the share of words the model gets wrong against a human reference—the standard measure of accuracy for live, streaming transcription.

Streaming word error rate on English audio.

*Lower is better*
Streaming word error rate (English)
AssemblyAI Universal-3.5 Pro Realtime
5.53%
Deepgram Flux
8.87%
Deepgram Nova-3
9.39%
Cartesia
10.10%

Source: AssemblyAI published benchmarks — assemblyai.com/benchmarks.

Common questions