Insights & Use Cases
September 8, 2026

Automatically determine video sections with AI using Python

Learn how to automatically split a video into sections in Python, generate catchy section titles with LLM Gateway, and format them as YouTube Chapters.

Ryan O'Connor
, 
Senior Developer Educator
Reviewed by
No items found.
Abstract green mobius illustration
Table of contents

YouTube Chapters are the closest thing video has to a table of contents. They turn a 25-minute recording into something a viewer can skim, and they turn a scrubber bar into a set of labelled destinations. The problem is that writing them is tedious — you have to watch the whole thing, decide where the topic shifts, and name each shift.

That is exactly the kind of work worth handing to a model.

In this tutorial we'll take a recorded meeting, split it into topic sections automatically, generate catchy titles for each one, and format the whole thing as a YouTube Chapters block you can paste straight into a video description. Everything runs in Python with two API calls.

Here's what we'll build, in order:

  1. Transcribe a video and get back topic sections with timestamps, using Speech Understanding Summarization.
  2. Convert those millisecond timestamps into the H:M:S format YouTube Chapters requires.
  3. Rewrite the section headlines into something a viewer would actually click, using LLM Gateway.
  4. Verify the model didn't quietly invent a timestamp.

The sample file is a GitLab logistics meeting — a real recording of people talking over each other about merge request rates, which is a much fairer test than a clean single-speaker podcast. The full code is in the rebuilt companion repo at github.com/kelsey-aai/automatic-video-sections.

A note on Auto Chapters

If you've built this before, you probably used the auto_chapters parameter. That parameter is deprecated, and Speech Understanding Summarization is what replaces it.

The good news is that the swap is small. Summarization returns the transcript broken into ordered topic sections, each with a headline, a summary and start and end timestamps — which is what chapters were. What changes is where you put the request and where you read the response. If you're migrating existing code, the Auto Chapters migration guide has a field-by-field mapping. If you're building this for the first time, just follow along below.

Getting started

First, a virtual environment.

# Mac/Linux:
python3 -m venv venv
. venv/bin/activate

# Windows:
python -m venv venv
venv\Scripts\activate.bat

We only need one dependency.

pip install requests

Now grab an API key from your AssemblyAI dashboard and put it in your environment. The free tier covers 185 hours of pre-recorded audio, which is far more than you'll need here.

# Mac/Linux:
export ASSEMBLYAI_API_KEY=<YOUR_KEY>

# Windows:
set ASSEMBLYAI_API_KEY=<YOUR_KEY>

Transcribing the meeting

One request does both jobs. We submit the video URL for transcription and, in the same body, ask for Summarization under the speech_understanding key.

import os
import time
import requests

BASE_URL = "https://api.assemblyai.com"
HEADERS = {"authorization": os.environ["ASSEMBLYAI_API_KEY"]}

AUDIO_URL = "https://storage.googleapis.com/aai-web-samples/meeting.mp4"

data = {
    "audio_url": AUDIO_URL,
    "speech_models": ["universal-3-5-pro"],
    "speech_understanding": {
        "request": {
            "summarization": {
                "summary_type": "paragraph"
            }
        }
    }
}

response = requests.post(f"{BASE_URL}/v2/transcript", json=data, headers=HEADERS)
response.raise_for_status()

transcript_id = response.json()["id"]
polling_endpoint = f"{BASE_URL}/v2/transcript/{transcript_id}"

while True:
    transcript = requests.get(polling_endpoint, headers=HEADERS).json()
    if transcript["status"] == "completed":
        break
    if transcript["status"] == "error":
        raise RuntimeError(f"Transcription failed: {transcript['error']}")
    time.sleep(3)

Two parameters are worth pausing on.

speech_models is a plural array on pre-recorded audio, and setting it explicitly matters more than it looks. Leave it out and your job runs on whatever your account default happens to be, which means the tutorial you copy today and the tutorial you copy in six months can produce different transcripts for reasons that have nothing to do with your code. We pin universal-3-5-pro. It handles overlapping speakers and cross-talk well, which a logistics meeting has plenty of, and it's the model our async benchmarks are measured on. Because we pin only the flagship, there is no automatic Universal-2 fallback here; add universal-2 to the array if you need coverage beyond its 18 languages. See selecting a speech model for the full picture.

summary_type takes paragraph or bullets. Paragraph gives you prose per section, which is what we want for descriptions. Bullets gives you something terser, which is better if you're rendering chapter markers into a cramped UI.

There's a third parameter you don't need yet but should know about: effort, which defaults to low. Bump it to medium when missed details actually cost you something — important meetings, multilingual audio, or files running past about 90 minutes. For a 25-minute standup, the default is fine.

Reading the sections

Now let's look at what came back.

# print the text
print(transcript["text"], end="\n\n")

summarization = transcript["speech_understanding"]["response"]["summarization"]

if summarization["status"] != "success":
    raise RuntimeError("Summarization did not complete successfully")

sections = summarization["summary"]

# now we print the video sections information
for section in sections:
    print(f"Start: {section['start']}, End: {section['end']}")
    print(f"Headline: {section['headline']}")
    print(f"Text: {section['text']}\n")

Three things to notice about the response shape.

The sections live at speech_understanding.response.summarization.summary, not at a top-level chapters field. Each element carries headline, text, start and end — so the old summary field is now text, and the old gist field is gone. If you were using gist for a short label, headline is what you want. And start and end are still milliseconds, which matters a lot in the next section.

Check status for success before you read the summaries. Transcription and Summarization can succeed and fail independently, and a missing check is the sort of thing that surfaces at 2am rather than in development.

Start Chaptering Your Own Video

Get an API key and run this tutorial on your own recordings. The free tier covers 185 hours of pre-recorded audio and 333 hours of streaming — no credit card needed.

Sign up free

Turning video sections into YouTube Chapters

We have sections. YouTube wants something specific.

YouTube Chapter timestamp formatting

YouTube's format is strict in ways that are easy to get wrong. Timestamps go at the start of the line, followed by a space and the title. They must be in ascending order. There have to be at least three of them. And — the one that catches everybody — the first chapter must start at zero. If your first section begins at 240ms because there was a beat of silence, YouTube ignores the entire block.

Our timestamps are in milliseconds, so first we need a converter.

def ms_to_hms(start):
    s, ms = divmod(start, 1000)
    m, s = divmod(s, 60)
    h, m = divmod(m, 60)
    return h, m, s

Then we build the lines. YouTube accepts either MM:SS or HH:MM:SS, so we look at the last section to decide which format the whole block uses — mixing them within one description is a good way to have YouTube silently drop your chapters.

def create_timestamps(sections):
    last_hour = ms_to_hms(sections[-1]["start"])[0]
    time_format = "{m:02d}:{s:02d}" if last_hour == 0 else "{h:02d}:{m:02d}:{s:02d}"

    lines = []
    for idx, section in enumerate(sections):
        # first YouTube timestamp must be at zero
        h, m, s = (0, 0, 0) if idx == 0 else ms_to_hms(section["start"])
        lines.append(f"{time_format.format(h=h, m=m, s=s)} {section['headline']}")

    return "\n".join(lines)
timestamp_lines = create_timestamps(sections)
print(timestamp_lines)

That's already a working chapters block. Paste it into a description and it works. But read the headlines and you'll see the problem — they're accurate and completely flat. "Data Team Lag Issues" describes the section. It doesn't make anyone click it.

Improving results with LLM Gateway

So let's rewrite them. LLM Gateway gives you one API across Anthropic, OpenAI, Google and Qwen models, with automatic cross-provider fallback — which means the rewriting step doesn't become a second vendor integration.

The prompt does the real work here. We give the model a role, the context that this is a GitLab logistics meeting, an explicit instruction about the input format, the timestamps themselves, and an exact output format.

prompt = f"""
ROLE:
You are a YouTube content professional. You are very competent and able to come up with catchy names for the different sections of video transcripts that are submitted to you.
CONTEXT:
This transcript is of a logistics meeting at GitLab
INSTRUCTION:
You are provided information about the sections of the transcript under TIMESTAMPS, where the format for each line is `<TIMESTAMP> <SECTION SUMMARY>`.
TIMESTAMPS:
{timestamp_lines}
FORMAT:
<TIMESTAMP> <CATCHY SECTION TITLE>
OUTPUT:
""".strip()

Then a single POST. The Gateway speaks the OpenAI chat-completions shape, so if you've called any modern LLM API this will look familiar — and your existing AssemblyAI key is the only credential involved.

llm_response = requests.post(
    "https://llm-gateway.assemblyai.com/v1/chat/completions",
    headers=HEADERS,
    json={
        "model": "claude-sonnet-4-6",
        "messages": [{"role": "user", "content": prompt}],
        "max_tokens": 1000,
    },
)
llm_response.raise_for_status()

# Extract the response text and print
output = llm_response.json()["choices"][0]["message"]["content"].strip()
print(output)

Swap model for any other ID on the available-models list if you want to compare — that's the point of a gateway. And log the request_id that comes back on every response; it's what support needs to find your exact call if something looks off.

The output is noticeably better:

00:00 Proposing Department Key Reviews
01:40 Clarifying Merge Request Rate Definitions
03:51 Fixing the Wider Merge Request Rate Metric
08:18 Confirming Wider Rate is Community-Only
08:40 Data Team Lag Issues
09:40 Discussing Postgres Replication
13:05 Defect Tracking and SLO Update
18:32 Smaller Security Metric Decline Seen as Improvement
20:42 Investigating Below-Target Narrow Merge Request Rate
23:50 Wrapping Up the Meeting

‍

Regex and verification

Here's the part most tutorials skip. The model was asked to preserve the timestamps exactly and rewrite only the titles. Usually it does. Sometimes it adds a preamble, or drops a line, or — worst case — nudges a timestamp.

If you're publishing this automatically, you cannot assume. So we verify.

First, strip anything that isn't a timestamped line.

import re

def filter_timestamps(text):
    lines = text.splitlines()
    # Use regex to filter lines starting with a timestamp
    timestamped_lines = [line for line in lines if re.match(r'\d+:\d+', line)]
    filtered_text = '\n'.join(timestamped_lines)
    return filtered_text

filtered_output = filter_timestamps(output)

Then check the model's output against what we generated ourselves. Two assertions, not one. First the line count, because zip stops at the shorter sequence — so if the model quietly drops a section, a naive loop compares the survivors, finds them all fine, and passes. That's the failure mode this whole step exists to catch. Then every individual timestamp. If either check fails we raise, rather than publishing a chapters block that points at the wrong moments.‍

original = timestamp_lines.splitlines()
filtered = filtered_output.splitlines()

if len(original) != len(filtered):
    raise RuntimeError(f"Line count mismatch - expected {len(original)}, got {len(filtered)}")

for o, f in zip(original, filtered):
    original_time = o.split(' ')[0]
    filtered_time = f.split(' ')[0]
    if not original_time == filtered_time:
        raise RuntimeError(f"Timestamp mismatch - original timestamp '{original_time}' does not match LLM timestamp '{filtered_time}'")

print(filtered_output)

‍

The timestamps came from the API. The titles came from the model. Keeping those two responsibilities separate — and asserting the boundary in code — is what makes this safe to run unattended.

See Voice AI In Action

Experience natural, real-time conversations that go far beyond IVR menus. Test streaming transcription speed and accuracy on your own audio.

Try playground

Where to take this next

You can publish these chapters automatically. YouTube's Data API lets you update a video's description programmatically once you've worked through server-side web app authentication. From there it's an update call on the video resource.

The same two-step pattern — structure from the transcription API, judgment from LLM Gateway — extends further than chapters. Add speaker labels and the same sections become an attributed meeting recap. Add action_items to the same speech_understanding request and you get structured follow-ups with quotes and timestamps alongside the chapters, in one call. Persist the sections with their millisecond offsets and you have the index for searchable transcripts across a whole video library.

Final words

The interesting thing about this build isn't the chaptering — it's the division of labour.

Summarization is deterministic about structure. It tells you where the topics change and gives you timestamps you can trust, because they come from the transcript rather than from a model's summary of one. The LLM never sees the audio and never decides where a section starts. It only renames things.

That's why the verification step at the end costs a dozen lines. Once the model's job is narrow enough, checking its work is trivial — and a pipeline you can check is a pipeline you can leave running. Most of the "AI wrote something wrong" failures in production content tooling aren't model failures at all. They're jobs that were scoped too broadly to verify.

Summarization is $0.03 per hour of audio on top of transcription, and the LLM Gateway call for a 25-minute meeting costs fractions of a cent. See pricing for the full breakdown.

Build Your Video Pipeline On Voice AI

Transcription, summarization, and LLM access behind one API key. Pay per second, no minimums, no upfront commits, unlimited concurrency.

Sign up free

Frequently asked questions

How to transcribe segments of a video with API calls?

Submit the video once and read back timestamped segments — you don't make one call per segment. A single POST to https://api.assemblyai.com/v2/transcript with a speech_understanding.request.summarization object returns the transcript split into ordered topic sections, each with headline, text, start and end in milliseconds, at speech_understanding.response.summarization.summary. If you need finer granularity than topic sections, the same transcript exposes per-word and per-sentence timestamps you can slice however you like.

Best API for video transcription and subtitles at scale

Look for per-second billing, unlimited concurrency, and no upfront commitment, because video libraries arrive in bursts rather than at a steady rate. AssemblyAI's pre-recorded transcription runs at $0.21/hr on Universal-3.5 Pro and $0.15/hr on Universal-2, billed per second with no minimums, and the same transcript exports to SRT and VTT at no extra charge. Speech Understanding features are priced à la carte on top — Summarization at $0.03 per hour of audio — so you only pay for the ones you turn on.

Most accurate API for transcribing YouTube and video content

Universal-3.5 Pro is our most accurate pre-recorded model, and it's what this tutorial pins with "speech_models": ["universal-3-5-pro"]. It covers 18 languages with native code-switching and falls back to Universal-2 automatically once you add universal-2 to that array, for 99 languages total. Our English async benchmarks put mean WER at 5.6% and median at 4.9% across 26 datasets and more than 80,000 files — but for video content the number that usually matters more is entity accuracy, since product names, people and companies are what viewers search for. Full methodology is on our benchmarks page.

How do I get started with AssemblyAI's LLM Gateway?

Send a POST to https://llm-gateway.assemblyai.com/v1/chat/completions with your existing AssemblyAI API key — there's no separate signup or second vendor account. The body takes model, messages, and max_tokens, following the OpenAI chat-completions shape, and you read the result from choices[0].message.content. One API covers Anthropic, OpenAI, Google and Qwen models with automatic cross-provider fallback, and an EU endpoint is available at llm-gateway.eu.assemblyai.com. The quickstart has the full parameter list.

How do I get started with AssemblyAI's API?

Create a free account, copy your API key from the dashboard, and send a POST to /v2/transcript with an audio_url — that's the whole first request. The free tier includes 185 hours of pre-recorded transcription and 333 hours of streaming, with no credit card required. From there you add capabilities as parameters on the same request rather than as separate integrations: speech_models to pin a model, speaker_labels for diarization, speech_understanding for summaries and action items.

What replaced Auto Chapters, and do I need to change my code?

Speech Understanding Summarization replaced the auto_chapters parameter, which is deprecated, so yes, existing code needs a change. Replace auto_chapters: true with a summarization object under speech_understanding.request, and read sections from speech_understanding.response.summarization.summary instead of the top-level chapters array. Per-section, summary becomes text and gist is gone; use headline for a short label. The migration guide maps every field.

‍

Title goes here

Lorem ipsum dolor sit amet, consectetur adipiscing elit, sed do eiusmod tempor incididunt ut labore et dolore magna aliqua. Ut enim ad minim veniam, quis nostrud exercitation ullamco laboris nisi ut aliquip ex ea commodo consequat. Duis aute irure dolor in reprehenderit in voluptate velit esse cillum dolore eu fugiat nulla pariatur.

Button Text
Tutorial
LLMs
Python