Summarize audio with LLMs in Node.js
Learn how to automatically summarize audio with Node.js and the AssemblyAI API.



Summarizing audio is a two-step job: convert speech to text, then pass that text to an LLM. In Node.js you can do both with one API key — AssemblyAI's speech-to-text API handles the transcription, and the LLM Gateway routes the transcript to Claude, Gemini, GPT, or a self-hosted open model without a second provider account. This guide builds a working script, then covers the parts that matter in production: picking the right model for your audio length, handling short clips in milliseconds, logging for support, and keeping data in a specific region.
What is audio summarization?
Audio summarization is the process of producing a short, structured description of what was said in a recording. It is almost never a single model call. The standard pipeline is transcription followed by summarization: a speech recognition model converts the audio to text, and a language model compresses that text into a paragraph, a bullet list, action items, or a JSON object.
That two-stage shape is the important thing to understand, because the two stages fail differently. The LLM stage fails visibly — you get a summary that misses the point, and you fix the prompt. The transcription stage fails invisibly. The LLM receives a clean, fluent, confidently wrong sentence and summarizes it faithfully.
How accurate is automated audio summarization?
Summary quality is bounded by transcript quality. An LLM cannot recover a name it never received. If the transcript renders a customer as "Anita Deshpande" when she said "Anita Deshmukh," every summary downstream carries the wrong name, and no amount of prompt engineering fixes it. The same holds for figures, product names, and speaker attribution — if the transcript assigns a commitment to the wrong speaker, the action-item list assigns it to the wrong person.
So the useful accuracy question is not "how good is the LLM at summarizing," it is "how much error is entering the pipeline before the LLM ever runs." Two measurements matter for that.
The first is word error rate on realistic audio. Multilingual and code-switched speech is where transcription models diverge most, because that is where they are most likely to substitute a plausible-sounding wrong word. On AssemblyAI's code-switching benchmark, measured as average normalized WER — lower is better:
The second is speaker attribution. Any summary that says "who agreed to what" depends on speaker diarization being right. Universal-3.5 Pro is optimized for cpWER — concatenated minimum-permutation word error rate, which penalizes both transcription errors and misattributed speech, rather than DER, which only scores segment boundaries. Average cpWER, lower is better:
Researchers score summarization itself with ROUGE against human-written references, and that is worth knowing about, but it is the wrong lever to pull first. Swapping summarization models moves quality at the margin. Cutting upstream transcription error moves it structurally. Full methodology for both tables is on the benchmarks page and in the Universal-3.5 Pro async writeup.
Set up your environment
You need Node.js 18 or higher and an AssemblyAI API key. Create a project:
mkdir audio-summarization
cd audio-summarization
npm init -y
Add ES module support to package.json:
{
"type": "module"
}
Install the SDK:
npm install --save assemblyai
Set your API key as an environment variable:
# macOS / Linux
export ASSEMBLYAI_API_KEY=<YOUR_KEY>
# Windows
set ASSEMBLYAI_API_KEY=<YOUR_KEY>
Transcribe and summarize audio with LLMs
The workflow is two calls. The SDK uploads and transcribes the audio, then you send the resulting transcript ID to the LLM Gateway. You never move the transcript text through your own process.
Transcribe the audio with Universal-3.5 Pro
Create summarize.js:
import { AssemblyAI } from 'assemblyai';
const client = new AssemblyAI({ apiKey: process.env.ASSEMBLYAI_API_KEY });
const audioUrl = 'https://assembly.ai/sports_injuries.mp3';
const transcript = await client.transcripts.transcribe({
audio: audioUrl,
speech_models: ['universal-3-5-pro']
});
if (transcript.status === 'error') {
throw new Error(`Transcription failed: ${transcript.error}`);
}
console.log(transcript.text);
speech_models is an array on the async API. Universal-3.5 Pro is the flagship model at $0.21/hr, with native code-switching across 18 languages. If you need coverage beyond those 18, add universal-2 — it spans 99 languages total at $0.15/hr, at a lower accuracy tier. Current rates for both are on the pricing page.
Summarize the transcript with LLM Gateway
Pass transcript_id and use the {{ transcript }} template tag in your prompt. The Gateway substitutes the transcript server-side, so you are not shipping the full text back up the wire:
const llmResponse = await fetch(
'https://llm-gateway.assemblyai.com/v1/chat/completions',
{
method: 'POST',
headers: {
authorization: process.env.ASSEMBLYAI_API_KEY,
'Content-Type': 'application/json'
},
body: JSON.stringify({
model: 'claude-sonnet-4-6',
messages: [
{
role: 'user',
content:
'Provide a concise, one-paragraph summary of this transcript.\n\n{{ transcript }}'
}
],
transcript_id: transcript.id,
max_tokens: 1000
})
}
);
const result = await llmResponse.json();
console.log(result.choices[0].message.content);
The spacing inside {{ transcript }} matters — the tag is matched literally. This is the current path for LLM-powered summarization and auto-chapters: Speech Understanding for the analysis layer, and the LLM Gateway for the generation. Full request options are in the chat completions reference.
Run the script
node summarize.js
Summarize short audio in milliseconds with the Sync API
If your audio is under 120 seconds — a voice note, a voicemail, a single meeting clip, a dictated command — you do not need the upload-and-poll flow at all. The Sync API takes the audio in one request and returns a transcript in the response. No upload step, no polling loop, no job to track.
import { AssemblyAI } from 'assemblyai';
const client = new AssemblyAI({ apiKey: process.env.ASSEMBLYAI_API_KEY });
const result = await client.sync.transcribe('./voice-note.wav', {
model: 'universal-3-5-pro'
});
console.log(result.text);
console.log('session_id:', result.session_id);
From there, summarizing is the same Gateway call — the only difference is that you pass the text directly instead of a transcript_id, because there is no async transcript object to reference:
const llmResponse = await fetch(
'https://llm-gateway.assemblyai.com/v1/chat/completions',
{
method: 'POST',
headers: {
authorization: process.env.ASSEMBLYAI_API_KEY,
'Content-Type': 'application/json'
},
body: JSON.stringify({
model: 'qwen3.5-4b-32k-fast',
messages: [
{
role: 'user',
content: `Summarize this voice note in one sentence.\n\n${result.text}`
}
],
max_tokens: 200
})
}
);
const summary = await llmResponse.json();
console.log(summary.choices[0].message.content);
In production logs the Sync API runs at roughly 127ms p50 and about 302ms on average. Paired with a latency-optimized summarization model, a full transcribe-and-summarize round trip for a short clip lands well inside a second.
Two constraints to plan around: Sync accepts WAV or raw PCM, and it does not ingest URLs, so remote files need downloading first. Audio over 120 seconds goes through the async flow above. Setup details are in the Sync API getting-started guide, and current rates are on the pricing page linked above.
Handle different audio formats and sources
The SDK accepts a URL, a local path, a buffer, or a stream. It detects which one you passed and handles the upload:
import fs from 'fs';
// Local file path
const fromPath = await client.transcripts.transcribe({
audio: './path/to/your/audio.mp3',
speech_models: ['universal-3-5-pro']
});
// Buffer
const audioBuffer = await fs.promises.readFile('./audio.mp3');
const fromBuffer = await client.transcripts.transcribe({
audio: audioBuffer,
speech_models: ['universal-3-5-pro']
});
// Readable stream
const audioStream = fs.createReadStream('./audio.mp3');
const fromStream = await client.transcripts.transcribe({
audio: audioStream,
speech_models: ['universal-3-5-pro']
});
Streams are the right choice when the audio is already moving through your application — an upload handler, a queue consumer — and you would rather not buffer the whole file in memory first.
Summarize a meeting recording
Generic prompts produce generic summaries. Giving the model the context it cannot infer from the audio — what the meeting was, who was in it, what you intend to do with the output — changes the result more than switching models does:
const meetingContext =
'This was a weekly sync for the Project Phoenix development team.';
const llmResponse = await fetch(
'https://llm-gateway.assemblyai.com/v1/chat/completions',
{
method: 'POST',
headers: {
authorization: process.env.ASSEMBLYAI_API_KEY,
'Content-Type': 'application/json'
},
body: JSON.stringify({
model: 'claude-sonnet-4-6',
messages: [
{
role: 'user',
content: `Context: ${meetingContext}
From the following meeting transcript, extract:
1. Key decisions made
2. Action items, each with an owner
3. Open questions or blockers
If an action item has no clear owner, say so rather than guessing.
{{ transcript }}`
}
],
transcript_id: transcript.id,
max_tokens: 1000
})
}
);
For multi-speaker recordings, enable speaker labels on the transcription request so the model can attribute statements. Owners in an action-item list are only as reliable as the diarization underneath them.
Customize summary type
The output shape is a prompt decision, not an API decision.
For a bulleted list, change only the prompt:
content:
'Generate a bulleted list of the main topics discussed. One bullet per topic, no preamble.\n\n{{ transcript }}'
For anything a downstream system will parse, do not ask for JSON in prose and hope. Use Structured Outputs with a response_format JSON schema so the response is constrained to the shape your code expects:
body: JSON.stringify({
model: 'claude-sonnet-4-6',
messages: [
{
role: 'user',
content: 'Summarize this transcript.\n\n{{ transcript }}'
}
],
transcript_id: transcript.id,
max_tokens: 1000,
response_format: {
type: 'json_schema',
json_schema: {
name: 'audio_summary',
schema: {
type: 'object',
properties: {
topic: { type: 'string' },
key_points: { type: 'array', items: { type: 'string' } },
action_items: { type: 'array', items: { type: 'string' } },
conclusion: { type: 'string' }
},
required: ['topic', 'key_points', 'action_items', 'conclusion']
}
}
}
})
Choose the right model for summarization
The Gateway exposes models from Anthropic, Google, OpenAI, and open-weight models AssemblyAI hosts directly, all behind one API key and one bill. For summarization the choice comes down to how long the audio is and how much nuance the summary has to preserve.
The practical rule: for a two-minute voice note there is no measurable summary-quality gain from a frontier model, and qwen3.5-4b-32k-fast costs and latencies an order of magnitude less. For a 90-minute multi-speaker call where the summary is the artifact people act on, use claude-sonnet-4-6. The LLM Gateway overview lists the full current lineup.
Switching between them is one string:
model: 'qwen3.5-4b-32k-fast' // was 'claude-sonnet-4-6'
Error handling and performance optimization
Log the request_id for every LLM call
Every Gateway response carries a request_id. Log it on success as well as failure — it is the first thing support will ask for, and it is unrecoverable after the fact:
try {
const transcript = await client.transcripts.transcribe({
audio: audioUrl,
speech_models: ['universal-3-5-pro']
});
if (transcript.status === 'error') {
throw new Error(`Transcription failed: ${transcript.error}`);
}
const llmResponse = await fetch(
'https://llm-gateway.assemblyai.com/v1/chat/completions',
{ /* request as above */ }
);
if (!llmResponse.ok) {
const errorData = await llmResponse.json();
throw new Error(`LLM Gateway request failed: ${JSON.stringify(errorData)}`);
}
const result = await llmResponse.json();
console.log('request_id:', result.request_id, 'model:', result.model);
} catch (error) {
console.error('An error occurred:', error.message);
}
Log result.model alongside it. When a fallback fires, that field is how you find out which model actually produced the summary.
Add fallback models for resilience
A fallback chain reroutes the request when the primary provider is degraded, without your code catching anything:
body: JSON.stringify({
model: 'claude-sonnet-4-6',
messages: [
{ role: 'user', content: 'Provide a concise summary.\n\n{{ transcript }}' }
],
transcript_id: transcript.id,
max_tokens: 1000,
fallbacks: [{ model: 'gemini-3.7-flash' }, { model: 'claude-haiku-4-5-20251001' }]
})
Use fallback_config with depth to control how far down the chain the Gateway will go. Order the chain by what a degraded summary is allowed to look like, not by price alone.
Use the EU endpoint for data residency
Point at the EU host and the request is processed in-region:
const EU_GATEWAY = 'https://llm-gateway.eu.assemblyai.com/v1/chat/completions';
Three things to factor into that decision. Regional routing carries a +10% surcharge over standard rates. The Gateway operates under Zero Data Retention — customer LLM responses are not stored. And the same residency option exists upstream for transcription at api.eu.assemblyai.com, at the same price as US, with audio and transcripts staying in the EU for GDPR purposes. If residency is a requirement, set it on both hops — an EU-only summarization call in front of a US transcription call does not give you EU residency.
Process multiple files in parallel
Transcription jobs are independent, so submit them together rather than in sequence:
const audioFiles = [
'https://storage.googleapis.com/aai-web-samples/espn.m4a',
'https://assembly.ai/sports_injuries.mp3'
];
const transcripts = await Promise.all(
audioFiles.map((audio) =>
client.transcripts.transcribe({
audio,
speech_models: ['universal-3-5-pro']
})
)
);
const completed = transcripts.filter((t) => t.status === 'completed');
const summaries = await Promise.all(
completed.map(async (t) => {
const res = await fetch(
'https://llm-gateway.assemblyai.com/v1/chat/completions',
{
method: 'POST',
headers: {
authorization: process.env.ASSEMBLYAI_API_KEY,
'Content-Type': 'application/json'
},
body: JSON.stringify({
model: 'claude-sonnet-4-6',
messages: [
{
role: 'user',
content: 'Provide a concise summary.\n\n{{ transcript }}'
}
],
transcript_id: t.id,
max_tokens: 1000
})
}
);
return res.json();
})
);
Filter on status === 'completed' before summarizing. A failed transcript still resolves its promise, and summarizing an empty transcript burns tokens to produce nothing.
What this looks like in production
Media and collaboration products lean on this pipeline hardest, because summaries are what make a library of recordings navigable. Veed, a browser-based video editor, put it this way:
"Assembly allowed our team to focus on what they are best at: Building a collaborative, browser-based video editor and distributing that product at speed and at velocity to our user base."
— Sabba Keynejad and Tim Mamedov, Veed
The pattern generalizes. The transcription and summarization layers are undifferentiated infrastructure; the product is what you build on top of the structured output.
Next steps
You now have the full path: transcribe with Universal-3.5 Pro, summarize through the LLM Gateway, constrain the output with a JSON schema, and route short clips through the Sync API instead of the async flow. The two decisions worth revisiting as you scale are model selection — most teams over-buy on the LLM and under-invest in transcription accuracy — and data residency, which is cheaper to get right before launch than after.
From here, look at Speech Understanding for chapters, entity detection, and sentiment alongside the summary, and revisit the Sync API section above if most of your audio is short.
Frequently asked questions
How do I summarize a local audio file instead of a URL?
Pass the file path directly to the audio parameter — the AssemblyAI Node.js SDK detects a local path and handles the upload for you. It also accepts a Buffer or a readable stream, which is useful when the audio is already flowing through your application and you would rather not write it to disk. The rest of the pipeline is unchanged: you get back a transcript object, and you pass its id to the LLM Gateway as transcript_id.
Can I change the format of the summary?
Yes. The output shape is controlled entirely by your prompt, so asking for a headline, a bullet list, or an action-item breakdown is a one-line change. When a downstream system has to parse the result, use Structured Outputs with a response_format JSON schema instead of asking for JSON in the prompt — that constrains the model to your schema rather than relying on it to comply.
Which LLM should I use to summarize audio?
It depends on audio length and how much nuance the summary must preserve. For short-form audio at volume, qwen3.5-4b-32k-fast is the cost and latency winner — self-hosted on AssemblyAI GPUs, a 32,768-token context window, averaging 612ms, at $0.10 per 1M input tokens and $0.50 per 1M output tokens. For long or nuanced recordings where the summary drives a decision, use claude-sonnet-4-6; gemini-3.7-flash is a reasonable mid-tier default between the two.
How do I summarize short audio quickly?
Use the Sync API. For audio under 120 seconds you can POST the file and get a transcript back in the same response — no upload step, no polling — which in production logs runs around 127ms at p50 and roughly 302ms on average. Pair it with a latency-optimized model like qwen3.5-4b-32k-fast and a full transcribe-and-summarize round trip for a voice note completes well under a second. Sync accepts WAV or raw PCM and does not ingest URLs, so download remote files first.
How accurate is LLM audio summarization?
Summary accuracy is capped by transcript accuracy, because an LLM cannot recover information that never reached it. Universal-3.5 Pro averages 7.69 normalized WER on AssemblyAI's code-switching benchmark, against 8.77 for ElevenLabs Scribe v2, 12.22 for Deepgram Nova-3 Multilingual, and 44.58 for OpenAI GPT-4o Transcribe, and averages 30.17 cpWER on diarization versus 35.26, 36.87, and 37.92 for Scribe v2, Gladia, and Nova-3 English. A misheard name or a misattributed commitment propagates into every summary downstream regardless of which LLM you run, so reducing upstream error moves quality more than changing summarization models does.
What is the best way to handle errors during transcription or summarization?
Check transcript.status before summarizing — a failed transcript resolves normally and carries an error field rather than throwing. Wrap the Gateway call in try...catch, check llmResponse.ok, and log the request_id and model from every response, not just failures, since request_id is what support needs to trace a request. Add a fallbacks chain so a degraded provider reroutes automatically instead of surfacing as an outage.
How can I optimize costs when processing large audio files?
Pass transcript_id with the {{ transcript }} tag rather than pasting transcript text into the prompt — the Gateway substitutes it server-side, so you are not paying to move the text twice. Then right-size the model: most short-form summarization does not benefit measurably from a frontier model, and qwen3.5-4b-32k-fast runs at a fraction of the cost. If you use regional routing for data residency, budget for the +10% surcharge on top of standard rates.
Can I process multiple audio files simultaneously?
Yes. Transcription jobs are independent, so submit them together with Promise.all() rather than awaiting each one in sequence, then summarize the completed transcripts in a second parallel pass. Filter on status === 'completed' between the two stages so failed jobs do not consume LLM tokens producing summaries of empty transcripts.
Lorem ipsum dolor sit amet, consectetur adipiscing elit, sed do eiusmod tempor incididunt ut labore et dolore magna aliqua. Ut enim ad minim veniam, quis nostrud exercitation ullamco laboris nisi ut aliquip ex ea commodo consequat. Duis aute irure dolor in reprehenderit in voluptate velit esse cillum dolore eu fugiat nulla pariatur.




