New Universal-3.6 Pro Realtime is now available Learn more
Insights & Use Cases

Async, sync, or realtime? How to choose the right speech-to-text API

A recap of AssemblyAI's October 2026 workshop on choosing between async, sync, realtime, and dictation speech-to-text APIs: a three-question framework, measured latency, pricing, and six mistakes to avoid.

Abstract green sphere illustration

Written by

Kelsey Foster

Published on

9 October 2026

"I'm building X. Which model should I use?"

That's the question Craig Bruder, a Forward Deployed Engineer on AssemblyAI's Onboarding team who joined in June, says he hears more than any other. It's not hard to see why. AssemblyAI now ships four distinct ways to turn speech into text — async, sync, realtime, and dictation — and new models have been landing at a pace of roughly two launches a month.

So we ran a live workshop to answer it directly. "Async, Sync, or Realtime: Which model do I choose?" was a free, interactive webinar AssemblyAI hosted live on Zoom on October 8, 2026, at 10am ET, with Q&A throughout. Craig Bruder presented, and Natalie Atkins from AssemblyAI's Events & Community team moderated. The session broke down the four main approaches to speech-to-text (async, sync, realtime, and dictation), and the recording is now on the AssemblyAI YouTube channel. It was built for developers and product teams picking an API for a voice agent, a live notetaker, captions, recorded-audio analytics, or voice input — and for anyone still figuring it out. A live poll at the start showed 29% of attendees were building voice agents, and voice input was one of the most popular answers.

This post recaps the framework, the numbers, and the audience Q&A, and adds verified specs and code from the docs so you can act on it today.

The short answer: pick the API based on when the audio exists and who is waiting for the text. If nobody's waiting, use async (pre-recorded). If someone's waiting and the audio is still being spoken, use realtime. If someone's waiting and the clip is already finished, use sync — or dictation if you want cleaned-up text back in the same call.

The workshop at a glance

Detail Answer
What Workshop: "Async, Sync, or Realtime: Which model do I choose?"
Who Hosted by AssemblyAI; presented by Craig Bruder (Forward Deployed Engineer, Onboarding); moderated by Natalie Atkins (Events & Community, Marketing)
Where Live webinar on Zoom; recording on the AssemblyAI YouTube channel
When October 8, 2026, 10am ET (about 30 minutes, including Q&A)
Why "Which model should I use?" is the most common question from builders, and the product line has grown fast
Who it's for Developers and product teams building voice agents, notetakers, captions, recorded-audio pipelines, or voice input
What you'll leave with A framework for matching latency, accuracy, and cost to the right API; when a hybrid setup makes sense; and the mistakes teams make as they scale

The framework: three questions that pick your API

Craig's core rule is simple: the shape of your latency decides the API, and price comes second. Work through three questions in order.

  1. Is someone waiting on the text right now? If no, you have a recorded file and latency doesn't matter much. Use async.
  2. If yes — is the person still talking? If the audio is live and you need words back as they're spoken, use realtime.
  3. If they've finished — do you want verbatim text or cleaned-up text? For a finished snippet you need transcribed fast, use sync. If you want the text cleaned and formatted in the same request, use dictation.

The questions Craig gets from customers map neatly onto this. "Should we transcribe live or after the call?" If nobody needs the text during the call, async is cheaper and more accurate because the model sees the whole file. "What's your P99 latency? We're building a voice agent." That's realtime. "Our old vendor's speaker labels lag about five seconds — too slow for a live call." That's a notetaker that needs live speaker labels, so realtime again.

And one attendee favorite: "What's the difference between realtime and streaming?" Trick question. They're the same thing.

If you're still comparing vendors rather than endpoints, start with our guide on how to choose the best speech-to-text API. This post assumes you've picked AssemblyAI and need to pick the right door.

The four speech-to-text APIs side by side

Async (pre-recorded) Sync Realtime (streaming) Dictation
How you call it Submit a job; get a webhook or poll One HTTP request, transcript in the response Open a WebSocket; stream audio, get turns back One HTTP request, verbatim + cleaned text back
Max audio Up to 10-hour files 80 ms to 2 minutes Sessions up to 3 hours Up to 2 minutes
Best for Call recordings, QA, coaching, meetings that already ended Short finished clips where you run your own turn detection Voice agents, live notetakers, live captions Voice input features that need send-ready text
Flagship model Universal-3.5 Pro Universal-3.5 Pro Universal-3.6 Pro Realtime Universal-3.5 Pro + hosted cleanup
Languages 18 (99 with Universal-2) 32 32 32
List price $0.21/hr $0.30/hr $0.45/hr base (billed on session time) $0.62/hr, cleanup included

Prices are pay-as-you-go list prices from the AssemblyAI pricing page, billed per second with no minimums.

Async: when the audio already exists and nobody's waiting

Async — the pre-recorded speech-to-text API — is the default for anything that's already recorded: call recordings, QA and coaching, interviews, a meeting that just ended. You submit a file (MP3, MP4, WAV, and most other common formats), and you get the transcript back via webhook or polling.

It's the cheapest option at $0.21/hr on Universal-3.5 Pro, it accepts files up to 10 hours, and it's typically the most accurate, because the model gets the whole conversation as context instead of a few seconds at a time. Add speaker diarization for +$0.02/hr and every utterance comes back labeled Speaker A, Speaker B, and so on.

Language coverage is the one place to check your requirements. Universal-3.5 Pro covers 18 languages with native code-switching. If you need more, set Universal-2 as a fallback and coverage extends to 99 languages. Craig also mentioned a new async model on the way that will match the 32-language realtime lineup.

Here's the first transcription from the pre-recorded quickstart in the docs (Python SDK):

import os

from assemblyai.prerecorded.v2 import Transcriber

transcriber = Transcriber(api_key=os.environ["ASSEMBLYAI_API_KEY"])

transcript = transcriber.transcribe("https://assembly.ai/wildfires.mp3")
print(transcript.text)

To label speakers, the docs pass a config with speaker_labels turned on:

from assemblyai.prerecorded.v2 import TranscriptionConfig

config = TranscriptionConfig(speaker_labels=True)
transcript = transcriber.transcribe("https://assembly.ai/wildfires.mp3", config=config)

for utterance in transcript.utterances:
    print(f"Speaker {utterance.speaker}: {utterance.text}")
Transcribe your first recording in minutes

Run async transcription with speaker labels on your own files. Start free with pay-as-you-go pricing and no commitments.

Sign up free

Sync: a finished clip, one request, no WebSocket

The Sync API is for short audio that's already finished — one utterance, one command, one turn. You send a single HTTP request and the transcript comes back in the response. No polling, no WebSocket, no session management.

The limits are tight by design: 80 milliseconds to 2 minutes of audio per request, WAV or raw PCM. Price is $0.30/hr.

import os

from assemblyai.sync.v1 import SyncTranscriber

transcriber = SyncTranscriber(api_key=os.environ["ASSEMBLYAI_API_KEY"])
transcriber.warm()  # open the connection ahead of the request
result = transcriber.transcribe("./sample.wav")
print(result.text)

An attendee asked the obvious follow-up: for a voice agent, is realtime the only choice, or does sync work for short turn-by-turn exchanges? Sync can work, with a trade-off. Realtime handles turn detection, utterance boundaries, and speaker labels for you across a session of up to three hours. With sync, you own the turn detection — the positioning line is "you detect the turn, we transcribe it." You skip the WebSocket complexity (dropped packets on a slow connection, session lifecycle), but you take on tracking turns yourself.

If you do chunk audio yourself, Craig's warning is the most important line in the workshop: chunk on utterances, not on fixed time windows. Cut a sentence in half every 5 or 30 seconds and word error rate climbs. He used a radiologist dictating findings on an X-ray as the example — a mid-sentence cut that garbles a term can have real downstream consequences. Use a voice activity detector (VAD) such as the open-source Silero VAD to find natural boundaries, then send each utterance to sync.

For a deeper side-by-side of these two recorded-audio paths, read sync vs. async transcription.

Realtime: when the person is still talking

Realtime — the streaming speech-to-text API — is for live audio. You open a WebSocket, stream audio in, and get transcripts back turn by turn, with end-of-turn detection built in. The flagship is Universal-3.6 Pro Realtime, launched September 29, 2026: 32 languages on one model with automatic language detection, at $0.45/hr base.

The question Craig says to ask here: does something talk back?

  • Yes — it's a voice agent, an automated phone line, anything that takes turns. Use Universal-3.6 Pro Realtime. You can tune the turn-detection settings (minimum and maximum delay before a turn ends), pass what the agent just said into the next turn as context, and turn on voice focus to suppress other voices near the microphone. Craig's example: a drive-thru agent shouldn't take the order from the kid in the back seat yelling for a Frosty.
  • No — it only listens. Meeting notetakers, live captions, and call recording need pure transcription and speaker labels, not agent-grade knobs. Craig previewed a lower-cost Universal-3.6 tier built for these listen-only use cases, coming soon. Until then, Universal-3.6 Pro Realtime handles both.

Fireflies, which runs AI meeting notes and a voice agent pipeline, put the selection criteria well when it evaluated the previous generation:

"We were searching for the best realtime ASR model for our voice agent pipeline in Fireflies. The new Universal 3.5 Pro speech model from Assembly is best so far in terms of accuracy, latency and language switching."

— Foysal Osmany, Software Engineer at Fireflies

Here's the streaming quickstart from the streaming docs. It transcribes a live radio stream, so you don't need a microphone to try it:

import os
import time

import requests
from assemblyai.streaming.v3 import (
    BeginEvent,
    Encoding,
    RealTimeError,
    RealTimeEvents,
    RealTimeParameters,
    RealTimeTranscriber,
    RealTimeTranscriberOptions,
    TerminationEvent,
    TurnEvent,
)

# A live AAC (ADTS) internet radio stream, so no microphone is needed.
STREAM_URL = "https://14123.live.streamtheworld.com/WBBRAMAAC.aac"
RUN_SECONDS = 25


def on_begin(client: RealTimeTranscriber, event: BeginEvent):
    print(f"Session started: {event.id}")
    print("Connected. Streaming live radio for ~25 seconds.")


def on_turn(client: RealTimeTranscriber, event: TurnEvent):
    if event.transcript:
        print(event.transcript)


def on_terminated(client: RealTimeTranscriber, event: TerminationEvent):
    print(f"Session terminated: {event.audio_duration_seconds}s of audio processed")


def on_error(client: RealTimeTranscriber, error: RealTimeError):
    print(f"Error: {error}")


def main():
    client = RealTimeTranscriber(
        RealTimeTranscriberOptions(terminate_timeout=30.0),
        api_key=os.environ["ASSEMBLYAI_API_KEY"],
    )

    client.on(RealTimeEvents.Begin, on_begin)
    client.on(RealTimeEvents.Turn, on_turn)
    client.on(RealTimeEvents.Termination, on_terminated)
    client.on(RealTimeEvents.Error, on_error)

    # AAC is self-describing (ADTS headers carry the sample rate)
    client.connect(
        RealTimeParameters(speech_model="universal-3-6-pro", encoding=Encoding.aac)
    )

    # Pull the live radio stream and forward each chunk to the transcriber.
    response = requests.get(STREAM_URL, stream=True)
    deadline = time.time() + RUN_SECONDS
    try:
        for chunk in response.iter_content(chunk_size=4096):
            client.stream(chunk)
            if time.time() > deadline:
                break
    finally:
        response.close()
        # Terminate finalizes the open turn.
        # Keep the connection open long enough to receive the last final.
        client.disconnect(terminate=True)


if __name__ == "__main__":
    main()

Notice the disconnect(terminate=True) in the finally block. That one line prevents the most expensive mistake on the list below.

Dictation: verbatim and cleaned-up text in one call

The first question from the audience: how is dictation different from streaming with realtime and cleaning up the text with an LLM yourself?

The answer is request count. The Dictation API is one endpoint that returns two things: the verbatim transcript and a cleaned-up version, ready to send. Craig's example: a user says, "So can you send me the Q3 numbers before the, uh, the Thursday meeting? No, wait, the Friday meeting. Thanks." You get back exactly what was said, plus a clean "Can you send me the Q3 numbers before the Friday meeting?" — no second request to an LLM, no token math.

Like sync, it's built for short clips (up to 2 minutes per request). It runs on Universal-3.5 Pro, costs $0.62/hr with the cleanup included, and supports a custom llm_instruction if you want your own house style instead of the default cleanup.

import assemblyai as aai

aai.settings.api_key = "<YOUR_API_KEY>"

result = aai.DictationTranscriber().transcribe_live("clip.wav")

print(result.text)  # verbatim transcript
print(result.llm_response)  # cleaned-up text, ready to send

If the rewrite ever fails, llm_response comes back null and you still get the verbatim text, so you can fall back gracefully. Craig noted strong demand from the medical space, where clinicians dictate notes in short bursts.

How fast is each API in practice?

Craig ran every endpoint himself two days before the workshop, three to seven runs each. These are round-trip numbers from a client, not server-side processing times, so treat them as a realistic floor for your own testing:

API Test Measured latency
Async 4-minute clip with speaker labels 10–32 seconds, submit to finished
Sync 27-second clip About 1 second round trip
Realtime Live stream First words about 1 second after connecting; final text a fraction of a second after the audio ends
Dictation Short clip About 1 second round trip, cleanup included

Two takeaways. First, the LLM cleanup in dictation adds very little latency compared to sync. Second, async latency scales with file length — it's fast, but it's never going to feel live. For realtime turn-taking specifically, Universal-3.6 Pro Realtime's median end-of-turn latency is 537 ms on AssemblyAI's 12,460-conversation English benchmark.

How AssemblyAI's models compare

Craig focused the benchmark section on async and realtime, because sync and dictation are served by the same flagship model family. The thread running through it: headline word error rate (WER) isn't the whole story.

Entities matter more than averages. A voice agent with a perfect WER on everything except the caller's email address still sends the confirmation to the wrong inbox. That's why AssemblyAI tracks missed entity rate — phone numbers, emails, locations, names — alongside WER. On the voice agent benchmark in the Universal-3.6 Pro Realtime launch post, the overall entity error rate is 14.4%, with names at 10.9%.

Independent benchmarks agree. Universal-3.6 Pro Realtime shows the lowest word error rate on Coval's rolling 7-day window (2.2%), and on Pipecat's open benchmark it posts a 0.96% pooled semantic WER. Semantic WER is worth understanding: if one model writes "Dr." and another writes "doctor," both got the meaning right, and semantic WER doesn't penalize either. See the full methodology on the AssemblyAI benchmarks page.

Universal-3.6 Pro vs. Universal-3.5 Pro Realtime, from the launch data:

  • Heavy background noise: WER drops from 7.08% to 6.13% — the birthday party, the busy car.
  • Short utterances: wrong confirmations fall from 2.65% to 1.45%, 45% fewer overall. "No" and "yep" matter a lot to a voice agent.
  • Code-switching: mean WER across five language pairs drops from 8.65% to 7.20%, including switches mid-turn.

And Craig's point about these numbers: they're out of the box. Prompting, keyterms, and turn-detection tuning can push results further on your own data.

Test all four APIs on your own audio

Compare async, sync, realtime, and dictation side by side in the browser. Upload a file or speak into your mic and see the latency and accuracy for yourself.

Try playground

Six mistakes teams make when choosing a speech-to-text API

This was the part of the workshop most attendees will want to bookmark.

1. Using async when someone is waiting

Teams pick async because it's the cheapest, then chop live audio into 5-second clips to fake a live experience. That's a round peg in a square hole, and the brutal mid-sentence cuts hurt accuracy. If you need fast results on your own chunks, run a VAD and send utterances to sync instead.

2. Streaming a recorded file through realtime

Realtime processes audio at the speed it's spoken. A 30-minute recording takes 30 minutes to finish — and you pay for the whole session. Send recordings to async.

3. Chopping long recordings to fit sync or dictation

Both cap at 2 minutes per request. If you must split a longer file, split on utterances, not fixed 30-second windows. Better yet, ask whether the job is really an async job.

4. Assuming every feature is on every API

Speaker labels and Medical Mode aren't on the Sync API yet. Check the docs for each endpoint before you design around a feature. If you build with Claude Code or another coding agent, connect it to the AssemblyAI docs MCP server so it picks up these per-endpoint differences automatically.

5. Picking the model by price

Start with the API that matches the shape of your product, then optimize cost from there. Picking by price first is how mistakes 1 and 3 happen.

6. Leaving WebSockets open

Realtime is billed on how long the session is open, not how much speech is in it. Leave a connection open and it runs until the 3-hour maximum — and you're billed for three hours. Send a Terminate message when the call ends, and set the inactivity_timeout connection parameter (5 to 3,600 seconds) as a backstop.

When a hybrid setup makes sense

The framework isn't either/or. A common production pattern is realtime for the live experience and async for the final record. A meeting notetaker streams live captions and speaker labels during the call, then re-runs the full recording through async afterward for the most accurate transcript, diarization, and summaries. Async sees the whole conversation as context, so the post-call transcript is the one you store. For more on that split, see our guide to real-time vs. batch transcription.

Another pairing from the session agenda: dictation for polished voice input running next to realtime for live captions. The same logic applies to voice products with both a live and a stored surface: an agent that talks to callers in realtime and a QA pipeline that reviews recorded calls in async are two different jobs, and each gets its own API.

More questions from the workshop Q&A

Is Arabic supported on the realtime model? Yes. Arabic is among the 32 languages on Universal-3.6 Pro Realtime.

Is the lower-cost listen-only realtime tier available yet? Craig said it's in preview and targeted for general availability later in October. Watch the AssemblyAI changelog for the official launch.

Does the realtime model handle French? Yes, French is one of the tier-one languages. Set it with a language parameter on the request, or let automatic language detection pick it up.

What about text-to-speech? AssemblyAI's Text-to-Speech API is launching later in October, with a dedicated preview webinar. Craig said to expect close to 100 voices at launch across languages and accents, with voice cloning on the roadmap. Parameter details, such as emotional tone controls, will come with the launch.

The real lesson: match the API to the moment

The workshop's framework sounds obvious once you hear it. But nearly every pitfall Craig listed — fake-streaming through async, streaming recorded files, chunking on the clock, sessions left open — comes from the same root cause: choosing by price or habit instead of by when the audio exists and who's waiting.

Here's the forward-looking part. As AssemblyAI's lineup grows — a listen-only realtime tier, a 32-language async model, text-to-speech — the number of doors will keep going up. The three questions won't change. Answer them first, prototype on the endpoint they point to, and optimize price last. You'll ship faster and spend less re-architecting later.

Watch the full recording above, and keep an eye on the AssemblyAI channel for the upcoming text-to-speech session.

Pick your API and start building

Async, sync, realtime, and dictation all run on one account with one API key. Get started free and make your first call in seconds.

Sign up free

Frequently asked questions

What's the difference between async, sync, and real-time speech-to-text?

Async (pre-recorded) transcription processes a complete recorded file in the background and returns the transcript by webhook or polling. Sync transcription takes a short, finished clip in one HTTP request and returns the transcript in the response. Real-time (streaming) transcription transcribes live audio over a WebSocket as people speak. The right choice depends on whether the audio already exists and whether someone is waiting for the text.

Is real-time transcription the same as streaming transcription?

Yes. "Real-time" and "streaming" speech-to-text describe the same thing: audio is sent continuously over a persistent connection and transcripts return turn by turn while the speaker is still talking. AssemblyAI's streaming API uses a WebSocket and is billed on session duration.

Should I transcribe a call live or after it ends?

Transcribe after the call with async if nobody needs the text during the call. Async is cheaper ($0.21/hr vs. $0.45/hr base for realtime) and typically more accurate, because the model sees the whole conversation. Use realtime only when a person or agent needs the words during the call, or run both: realtime for the live experience and async for the final record.

Can I use a synchronous speech-to-text API for a voice agent?

Yes, if you handle turn detection yourself. AssemblyAI's Sync API transcribes clips from 80 ms up to 2 minutes in a single HTTP request, so you can run a voice activity detector, cut audio at utterance boundaries, and send each turn. If you'd rather not manage turns, use a realtime model like Universal-3.6 Pro Realtime, which detects turns and passes context between them for you.

What is a dictation API, and how is it different from speech-to-text?

A dictation API returns both a verbatim transcript and a cleaned-up version of what the user meant, removing fillers and self-corrections. AssemblyAI's Dictation API does this in one request for clips up to 2 minutes, at $0.62/hr with cleanup included. Standard speech-to-text returns only the verbatim transcript, so you'd need a second LLM call to get the same result.

How much does each AssemblyAI speech-to-text API cost?

At pay-as-you-go list prices, async on Universal-3.5 Pro is $0.21/hr, the Sync API is $0.30/hr, Universal-3.6 Pro Realtime is $0.45/hr base, and the Dictation API is $0.62/hr. All are billed per second with no minimums, and add-ons such as speaker diarization are priced separately. Realtime is billed on how long the session stays open, so always terminate sessions when the audio ends.