Insights & Use Cases
September 1, 2026

Can transcripts be used to generate meeting agendas?

Meeting notes help you capture key decisions, action items, and important discussions so your team stays organized and productive after every meeting.

Kelsey Foster
, 
Growth
Reviewed by
No items found.
Abstract green cylinder illustration
Table of contents

Most AI notetakers fail in the same place. Not at summarization — at the transcript underneath it.

A summary written on top of a transcript that missed who said what, dropped the two-word interruption where someone actually agreed to the deadline, or garbled a product name into a homophone will be confidently wrong. And confidently wrong meeting notes are worse than no notes, because people stop reading them and start rewatching the recording.

So this post works bottom-up. What the transcript layer has to get right, how to feed the model meeting context so it gets those things right, what to build summaries and action items on now that the old summarization path is moving, and how to turn all of it into a next-meeting agenda — the original question this page was written to answer.

What does an AI notetaker actually need from its transcript?

Three things: correct words, correct speaker attribution, and timestamps you can jump to. Summaries, action items, and agendas are all derived products. Every one of them inherits whatever the transcript got wrong.

Meeting audio is the hard case for all three. People talk over each other. Turns are short — "yeah," "wait, no," "by Friday?" — and short turns are exactly what a diarization system tends to smear into the neighboring speaker. Someone joins from a car. Someone's laptop mic is on the far side of the table. A product name that exists nowhere in the training data comes up nine times.

None of that is a summarization problem. It's a speech-to-text problem, and it's where the quality of your notetaker is decided.

Why speaker accuracy decides whether meeting notes are usable

An action item without a correctly attributed owner isn't an action item. It's a sentence.

This is why the diarization metric matters more than the headline word error rate for this use case. Diarization error rate measures speaker segmentation on its own, in isolation from the words. Concatenated minimum-permutation word error rate — cpWER — measures what you'd actually ship: the words and the speaker labels together, scored as one output. If the model gets a word right but hands it to the wrong person, cpWER counts that as the error it is.

Universal-3.5 Pro produces the transcript and the speaker changes jointly rather than running diarization as a second pass over finished text, and it's optimized for cpWER rather than DER. That joint approach is what catches the short turn, the rapid back-and-forth, and the overlapped speech that a bolt-on speaker pass tends to lose.

Model cpWER (lower is better)
Universal-3.5 Pro 30.17
ElevenLabs Scribe v2 35.26
Gladia 36.87
Deepgram Nova-3 English 37.92

Source: AssemblyAI benchmarks.

Turn speaker diarization on with speaker_labels: true on async. It's a +$0.02/hr add-on on top of the $0.21/hr flagship rate.

Speed matters too, and not for the reason people usually give. It isn't about a number on a dashboard — it's about whether the notes are there when the participant closes the tab.

"The speed difference is immediately noticeable — our users see their conversations transcribed almost instantaneously. It feels so much more responsive than what we were using before."

— Jonathan Kim, Software Engineer, Granola

Start with the transcript layer

Run your own meeting audio through Universal-3.5 Pro with speaker labels on and see how the crosstalk comes through. No credit card, no sales call.

Sign up free

How do you give the model meeting context before it transcribes?

Prime it. Contextual prompting lets you hand Universal-3.5 Pro the things it couldn't possibly know from the audio alone, before the audio is processed.

For a meeting, that's a short, specific list:

  • The agenda or calendar title, which tells the model what the conversation is about
  • Participant names, so "Siobhan" doesn't come out as "Shivon"
  • Your company and product names, competitor names, and internal acronyms
  • The organization or domain the meeting belongs to

This is doing real work, not decoration. Recruiting platform Metaview built exactly this pattern — piping calendar titles, organizations, domains, and participant names into transcription — because interview transcripts live or die on getting candidate names, employer names, and role-specific vocabulary right.

The same idea shows up elsewhere with numbers attached: in healthcare testing, feeding a patient's prior-visit note into the model cut missed medical terms by 31%. Meetings have the same shape of problem. The vocabulary that matters most is the vocabulary that's specific to your customer.

Keep the prompt tight. A focused list of names and terms beats a wall of background text.

Which summarization path should you build on now?

Not the old one. If your notetaker calls summarization, summary_type, or auto_chapters, those parameters work on Universal-2 only, and only until September 15, 2026. This one will break for readers who copy an older integration, so check before you ship.

Here's where they go:

  • Summaries and action items move to the speech_understanding request object, alongside the rest of the speech understanding feature set — entity detection, topic detection, sentiment, key phrases, PII redaction.
  • Chapters move to LLM Gateway, which gives you one API into OpenAI, Anthropic, Google, and others. The LLM Gateway relaunch covers what changed and why.

The migration is worth doing on its own merits. Bullet-point notes, decisions, owner-tagged action items, and a next-meeting agenda are four different prompts over the same diarized transcript, and running them through a gateway means you pick the model per job and swap it later without rewriting your integration. There's more on the general pattern in our guide to summarizing audio at scale.

One habit worth building in from the start: pass the transcript with speaker labels and timestamps intact, not as flat text. A summarizer that can see who said what produces action items with owners. One that can't produces a paragraph.

Can transcripts be used to generate meeting agendas?

Yes — and it's the highest-value thing you can do with a meeting archive, because an agenda built from the last meeting is the only artifact that changes what happens in the next one.

The pattern is straightforward once the transcript layer is solid. Take the diarized transcript from the previous session, or several sessions in a recurring series, and extract four things:

  • Open action items — commitments with an owner and no recorded completion
  • Unresolved threads — topics raised that ended without a decision
  • Decisions made — so the next meeting doesn't relitigate them
  • Deferred items — anything explicitly pushed to "next time"

Order those by whether they're blocking someone, and you have a draft agenda. Timestamps make it defensible: every line can link straight back to the moment in the recording where it was said, which is what turns an agenda from a suggestion into something people trust.

Then keep a human in the loop. The model is good at recall and bad at knowing what the team decided offline in Slack yesterday. Generated agenda, human edit, sent — that's the workflow that survives contact with an actual team.

Prototype the agenda prompt in the playground

Upload a real meeting recording, turn on speaker labels, and try your summary and agenda prompts against the transcript before you write any integration code.

Try playground

How does this connect to the meeting platforms you already use?

Through the recording, not through a per-platform integration you maintain forever.

Meeting bot infrastructure like Recall.ai joins calls across the major conferencing platforms and hands you the audio. You transcribe that audio through the API and never write a platform-specific code path. Alternatively, if you're already capturing recordings — a call recorder, a contact center platform, a mobile app — you're one upload away.

Downstream is a webhook. When the transcript completes, push it to your own service, and from there into a CRM, a project tracker, a data warehouse, or a BI tool. Because LLM Gateway and speech understanding return structured JSON rather than prose, the fields drop into an analytics schema without a parsing layer in between: entities, topics, sentiment, speaker turns, timestamps, and your generated summary as separate columns.

Teams building live meeting experiences — in-call prompts, real-time coaching — run the same pipeline over the streaming endpoint instead.

"We were searching for the best realtime ASR model for our voice agent pipeline in Fireflies. The new Universal 3.5 Pro speech model from Assembly is best so far in terms of accuracy, latency and language switching."

— Foysal Osmany, Software Engineer at Fireflies

What about compliance for meeting data?

Meetings pick up everything — customer names, contract terms, salary figures, patient details in a clinical setting. So the controls matter as much as the accuracy.

Control What it covers
SOC 2 Type 2 · BAA available · GDPR Standard Business Associate Addendum for customers processing PHI, signable in minutes without a sales call.
PII redaction Applied across both transcript text and the audio file itself.
EU data residency api.eu.assemblyai.com, same price as US, data stays in the EU.
Self-hosted deployment Runs in your own cloud when the data can't leave it.

On health data specifically, here's the exact position: AssemblyAI enables covered entities and their business associates subject to HIPAA to use the AssemblyAI services to process protected health information (PHI). AssemblyAI is considered a business associate under HIPAA, and we offer a standard Business Associate Addendum (BAA) that is required under HIPAA to ensure that AssemblyAI appropriately safeguards PHI. The BAA can be signed in minutes, without a sales call. The full configuration reference lives in the API documentation.

The part most teams get backwards

Teams building notetakers usually spend their first month on prompt engineering and their second month discovering that the prompt was never the problem.

Flip it. Get the diarized transcript right — joint transcription and speaker attribution, primed with the agenda and the participant names, stored in your own database — and the summary prompt becomes almost boring. Which is the goal. The interesting work in an AI notetaker isn't the model call at the end. It's everything that makes the model call trivially easy.

You can see the full stack for this on our AI notetaker solutions page, and the per-second rates, including the diarization and speech understanding add-ons, on pricing.

Build your notetaker on Universal-3.5 Pro

Pay per second, no minimums, and speaker attribution built into the model rather than bolted on after it.

Sign up free

Frequently asked questions

What is the best speech-to-text API to build AI notetakers?

For notetakers, pick the API with the strongest joint transcription and speaker attribution, because attribution is what makes action items usable. Universal-3.5 Pro produces the transcript and the speaker changes together and is optimized for cpWER at 30.17 average, ahead of Deepgram Nova-3 English at 37.92, ElevenLabs Scribe v2 at 35.26, and Gladia at 36.87. Also weigh turnaround speed and whether the provider offers summarization, entity detection, and redaction on the same platform.

Can AssemblyAI integrate with existing meeting platforms?

Yes, usually through the recording rather than through a platform-specific integration. Meeting bot providers such as Recall.ai join calls across major conferencing platforms and hand you the audio, which you then submit to the API like any other file. If you already capture recordings yourself, you can upload them directly and receive results via webhook.

Is there an API to create bullet-point summaries from transcripts?

Yes. Summaries and action items are moving to the speech_understanding request object, and chapters move to LLM Gateway, where you choose the model and control the output format directly in your prompt. The older summarization, summary_type, and auto_chapters parameters run on Universal-2 only and are supported until September 15, 2026, so new builds should start on the new path.

How do I create an AI meeting scribe that generates summaries and action items?

Submit the recording with speaker_labels: true and a contextual prompt containing the agenda, participant names, and your product vocabulary, then pass the resulting diarized transcript, with timestamps intact, to an LLM prompt that asks for decisions, owner-tagged action items, and open threads. Keeping speaker labels in the text handed to the model is what makes owners come out correct rather than generic. Persist the transcript to your own storage, then render the notes from it.

How do I integrate transcript analysis with business analytics tools?

Take the webhook payload and write it to your warehouse as structured rows rather than storing rendered notes as text. Speech understanding returns entities, topics, sentiment, and key phrases as JSON fields alongside speaker turns and word-level timestamps, so they map cleanly onto columns that BI tools can group and filter. From there, meeting data joins to CRM records the same way any other event stream does.

Is AssemblyAI HIPAA-compliant?

AssemblyAI offers a standard Business Associate Addendum (BAA), which is what covered entities and their business associates subject to HIPAA need in place with a vendor processing protected health information. AssemblyAI is considered a business associate under HIPAA, and the BAA can be signed in minutes without a sales call. Alongside that, the platform carries SOC 2 Type 2, offers PII and PHI redaction across audio and transcripts, and supports EU data residency and self-hosted deployment.

‍

Title goes here

Lorem ipsum dolor sit amet, consectetur adipiscing elit, sed do eiusmod tempor incididunt ut labore et dolore magna aliqua. Ut enim ad minim veniam, quis nostrud exercitation ullamco laboris nisi ut aliquip ex ea commodo consequat. Duis aute irure dolor in reprehenderit in voluptate velit esse cillum dolore eu fugiat nulla pariatur.

Button Text
Virtual Meetings
Speech-to-Text