Every subtitle file is doing the same small job: pairing a piece of text with a window of time. That's it. The reason there are dozens of formats for a job that simple is that the people who invented them were solving very different problems — DVD authoring, broadcast closed captioning, fansubbing anime, and eventually the open web.
The practical result is that you'll hit a wall at the worst possible moment. You export captions, upload them, and YouTube accepts them but your HTML5 player silently shows nothing. Or your editor imports the file and every cue is off by three hours because of a timestamp separator.
So this is the reference: what each of the three formats you'll actually encounter contains, how their syntax differs line by line, which one to reach for in which situation, and how to generate all three from a single transcript. No tool roundup — just the formats.
What is an SRT file?
An SRT file (SubRip Subtitle, .srt) is a plain-text file that holds numbered subtitle blocks, each with a start time, an end time, and one or two lines of text. It came out of SubRip, a Windows program from the early 2000s that ripped subtitles off DVDs, and it stuck around because it's the simplest thing that could possibly work.
A block looks like this:
1
00:00:00,000 --> 00:00:02,880
Hi, I'm here to talk about the future of
2
00:00:02,880 --> 00:00:05,760
speech recognition and what it means for
developers building products today.Four parts, in this exact order:
- A sequence number. Integers starting at 1, incrementing by one. Most players tolerate gaps, but don't test it.
- A timecode line. HH:MM:SS,mmm --> HH:MM:SS,mmm. Hours are two digits, milliseconds are three, and the separator between seconds and milliseconds is a comma. The arrow is two hyphens and a right angle bracket, with a space on each side.
- One or more lines of text. Two lines is the convention; three starts to crowd the frame.
- A blank line to close the block.
That comma is the single most common source of broken SRT files. WebVTT uses a period in the same position, and a file that mixes the two will fail to parse in a strict player without telling you which line went wrong.
A few more things worth knowing before you generate SRT at scale:
- Encoding matters. SRT has no encoding declaration, so a player has to guess. UTF-8 without a byte order mark is the safe default; UTF-8 with a BOM will put a stray character in front of your first subtitle in some older desktop players.
- Styling is basically absent. The spec has nothing to say about position, color, or font. Some players honor a small subset of HTML tags inside the text — <i>, <b>, , and sometimes <font color="#ffffff"> — but it's a convention, not a guarantee. Don't build a design around it.
- Timestamps are absolute, not frame-based. SRT knows nothing about frame rate. This is usually a feature, but it means a file cut for 23.976fps content doesn't automatically conform if the video is later pulled up to 24fps.
- There's no metadata. No language tag, no speaker labels, no title. The filename is the only place to encode a language, which is why you see video.en.srt and video.es.srt conventions everywhere.
SRT is the format you default to when you don't know what the destination supports, because essentially everything supports it. YouTube, Vimeo, Facebook, LinkedIn, Premiere Pro, DaVinci Resolve, Final Cut, VLC, Plex — all of them read SRT.
What is a VTT file?
A WebVTT file (.vtt) is the W3C-standardized caption format built for HTML5 video. It does everything SRT does, and then adds the things SRT can't do: positioning, styling, cue metadata, region definitions, and chapter markers.
The same two cues in WebVTT:
WEBVTT
Kind: captions
Language: en
00:00:00.000 --> 00:00:02.880
Hi, I'm here to talk about the future of
00:00:02.880 --> 00:00:05.760
speech recognition and what it means for
developers building products today.What changed from the SRT version:
- The WEBVTT header is mandatory. The file must begin with it. A VTT file without that first line isn't a VTT file, and browsers will reject it outright.
- Milliseconds are separated by a period, not a comma. This is the difference that bites people converting between formats with a naive find-and-replace.
- Cue identifiers are optional. You can label a cue with any string on the line before the timestamp — intro, speaker-1, 42 — and address it from CSS or JavaScript. SRT's mandatory sequence numbers become optional, meaningful identifiers.
- Metadata headers are allowed right after WEBVTT: Kind, Language, and arbitrary key-value pairs. That language tag is a real advantage over SRT's filename-convention approach.
- Hours can be omitted for content under an hour: 00:02.880 is valid. Include them anyway; it costs nothing and removes ambiguity.
Then there's the part that makes VTT worth choosing. Cue settings go on the timestamp line, after the end time:
00:00:05.760 --> 00:00:09.200 line:90% position:50% align:center size:80%
Captions can be positioned anywhere in the frame.
00:00:09.200 --> 00:00:12.400 vertical:rl line:0
And they can run vertically for Japanese or Chinese text.And you can carry styling in the file itself, using a STYLE block with real CSS:
WEBVTT
STYLE
::cue(.speaker-a) {
color: #7cd6ff;
font-style: italic;
}
00:00:00.000 --> 00:00:02.880
<c.speaker-a>So where does this leave the format question?</c>WebVTT also supports NOTE blocks for comments that never render, REGION definitions for scrolling roll-up captions, and a Kind: chapters file type that turns the same syntax into chapter markers for a player's scrubber.
The catch: everything past the basic cue structure is inconsistently implemented. Positioning and ::cue styling work well in modern browsers using the native <track> element. Upload the same file to a social platform and the styling is usually discarded on ingest. Treat the advanced features as an enhancement for players you control, not as something the file guarantees.
What is a TXT transcript?
A .txt file is the transcript without the timing — just the words, usually with paragraph breaks and sometimes with speaker labels.
Speaker A: Hi, I'm here to talk about the future of speech recognition
and what it means for developers building products today.
Speaker B: Let's start with accuracy. Where are we actually at?This isn't a subtitle format, and it will not work as one. Every player that accepts a caption file needs timecodes; a TXT file has none. But it's the output people want far more often than they expect to, because most of what a transcript gets used for has nothing to do with playback:
- Feeding a summarizer, a classifier, or a search index
- Blog posts, show notes, and SEO-indexable page content
- Compliance archives and legal discovery
- Quoting, translating, or repurposing the content
If you want the words plus who said them, speaker labels come from speaker diarization, which segments the audio by speaker and tags each utterance. That turns a wall of text into something a human can actually read.
There's a middle option worth knowing about too: a timestamped transcript, where each paragraph carries a start time but the file isn't structured as cues. It's not a caption format, but it's what you want behind an interactive transcript widget where clicking a sentence seeks the player.
SRT vs VTT vs TXT: a direct comparison
| Capability | SRT | VTT (WebVTT) | TXT |
|---|---|---|---|
| Timing | Yes — HH:MM:SS,mmm (comma) | Yes — HH:MM:SS.mmm (period), hours optional | No |
| Styling | Informal only — some players honor <i>, <b>, <font> | Yes — STYLE blocks with real CSS and ::cue selectors | No |
| Positioning | No | Yes — line, position, align, size, vertical, regions | No |
| Metadata | None — language lives in the filename | Yes — Kind, Language, cue IDs, NOTE comments | Speaker labels only, by convention |
| Browser support | Not supported by <track> — convert first | Native in every modern browser via <track> | Renders as text, never as captions |
| Platform support | Near-universal — social, NLEs, desktop players | Broad on web and streaming; patchier in editing suites | Universal as a document, irrelevant as captions |
| Typical use | Social uploads, video editors, the safe default | HTML5 players, HLS/DASH streaming, chapters | Summarization, search, archives, repurposing |
If you only remember one thing: SRT for upload, VTT for the web, TXT for everything downstream of playback.
Transcribe a file and export SRT, VTT, or plain text from the same result. Accurate timestamps, clear docs, and a free API key to start.
How to create an SRT file from a transcript
The mechanics are the same whichever speech-to-text API you use: transcribe the audio, then ask for the subtitle export rather than the JSON.
With the AssemblyAI API, that's a dedicated endpoint on the finished transcript. You transcribe once and export as many formats as you need — the timings are already in the result, so the export is free of an extra transcription pass.
Python
import assemblyai as aai
aai.settings.api_key = "YOUR_API_KEY"
transcriber = aai.Transcriber()
config = aai.TranscriptionConfig(speech_models=["universal-3-5-pro"])
transcript = transcriber.transcribe("https://example.com/interview.mp4", config)
# SRT, with each caption capped at 32 characters per line
srt = transcript.export_subtitles_srt(chars_per_caption=32)
with open("interview.srt", "w", encoding="utf-8") as f:
f.write(srt)
# WebVTT from the same transcript
vtt = transcript.export_subtitles_vtt(chars_per_caption=32)
with open("interview.vtt", "w", encoding="utf-8") as f:
f.write(vtt)
# Plain text
with open("interview.txt", "w", encoding="utf-8") as f:
f.write(transcript.text)JavaScript
import { AssemblyAI } from "assemblyai";
import { writeFileSync } from "node:fs";
const client = new AssemblyAI({ apiKey: process.env.ASSEMBLYAI_API_KEY });
const transcript = await client.transcripts.transcribe({
audio_url: "https://example.com/interview.mp4",
speech_models: ["universal-3-5-pro"],
});
const srt = await client.transcripts.subtitles(transcript.id, "srt", 32);
writeFileSync("interview.srt", srt, "utf-8");
const vtt = await client.transcripts.subtitles(transcript.id, "vtt", 32);
writeFileSync("interview.vtt", vtt, "utf-8");
writeFileSync("interview.txt", transcript.text ?? "", "utf-8");curl
curl "https://api.assemblyai.com/v2/transcript/${TRANSCRIPT_ID}/srt?chars_per_caption=32" \
-H "authorization: ${ASSEMBLYAI_API_KEY}" \
-o interview.srt
curl "https://api.assemblyai.com/v2/transcript/${TRANSCRIPT_ID}/vtt?chars_per_caption=32" \
-H "authorization: ${ASSEMBLYAI_API_KEY}" \
-o interview.vttThat chars_per_caption parameter is the one knob most people ignore and shouldn't. It controls how many characters land on a caption line before the exporter breaks to a new one. Leave it unset and you get whatever the natural utterance boundaries produce, which on a fast speaker can be a wall of text that covers a third of the frame. For a 16:9 video, 32 to 42 characters per line is the range that reads comfortably; vertical video wants closer to 20.
Everything above runs on Universal-3.5 Pro, which matters for captions more than people expect. Caption timing is only as good as the word-level timestamps underneath it, and word timings are only as good as the transcript. A model that drops a word doesn't just lose the word — it shifts the cue boundary around it.
Beyond SRT and VTT: the other caption files
You'll run into these less often, but it's worth knowing what they are when a broadcaster or a localization vendor asks for one.
| Format | Extension | What it's for |
|---|---|---|
| SubStation Alpha / Advanced SSA | .ssa / .ass | Heavy typesetting, karaoke effects, animation. The fansub standard. |
| Timed Text Markup Language | .ttml / .dfxp / .xml | XML-based. Used by Netflix, broadcast workflows, and IMSC profiles. |
| Scenarist Closed Captions | .scc | Frame-accurate CEA-608 broadcast captions. Frame rate is part of the file. |
| SAMI | .smi | Legacy Microsoft HTML-like format. You'll only meet it in old archives. |
| SubViewer | .sub | Old DVD-era format, sometimes paired with an .idx index. |
The pattern here is that the exotic formats exist because of a delivery constraint — a broadcaster's frame-accurate ingest, a localization house's XML pipeline, a fansub group's typesetting. If nobody is imposing that constraint on you, SRT and VTT cover the ground.
Which format should you use where?
Uploading to a social platform. SRT. YouTube, LinkedIn, Facebook, and Instagram all accept it, and the styling you'd gain from VTT gets stripped on ingest anyway. YouTube also accepts VTT and will honor basic positioning, but SRT is one less thing to debug.
Your own HTML5 player. VTT, no contest. The <track> element only speaks WebVTT:
<video controls src="/media/interview.mp4">
<track kind="captions" src="/media/interview.en.vtt" srclang="en" label="English" default>
<track kind="captions" src="/media/interview.es.vtt" srclang="es" label="Español">
<track kind="chapters" src="/media/interview.chapters.vtt" srclang="en">
</video>Hand that same element an SRT file and you get silence — no error, no captions, nothing in the console. It's one of the most quietly confusing failure modes in web video.
HLS or DASH streaming. VTT, segmented. Both adaptive protocols carry WebVTT as a subtitle rendition, which means the player can switch languages without reloading the video.
A video editor. SRT for import, always. Premiere, Resolve, and Final Cut all read it, and all of them prefer to apply their own styling anyway. You're handing the editor timing and words; they handle the look.
Broadcast delivery. Whatever the spec sheet says, and it will say SCC or TTML. Get the frame rate in writing before you generate anything.
Anything that isn't playback. TXT. If the words are going into a summarizer, a search index, a compliance archive, or a blog post, timing is dead weight.
Live captions are a different problem
Everything so far assumes a finished file. Live captioning — a webinar, a broadcast, a support call, an accessibility overlay — works differently, because there's no file to export. Text arrives as the audio does.
With streaming speech-to-text, you hold a WebSocket open and receive partial transcripts that update in place, then a final transcript when the model decides the turn is complete. Your renderer paints the partials into a caption bar and replaces them as the finals land.
import os
import time
from assemblyai.streaming.v3 import (
RealTimeEvents,
RealTimeParameters,
RealTimeTranscriber,
TurnEvent,
)
client = RealTimeTranscriber(api_key=os.environ["ASSEMBLYAI_API_KEY"])
def on_turn(client: RealTimeTranscriber, event: TurnEvent) -> None:
if not event.transcript:
return
if event.end_of_turn:
print(event.transcript) # commit this cue
else:
print(event.transcript, end="\r") # repaint the live line
client.on(RealTimeEvents.Turn, on_turn)
client.connect(RealTimeParameters(speech_model="universal-3-5-pro", sample_rate=16000))
# 3200 bytes is 100 ms of 16 kHz mono PCM16. The streaming API accepts between
# 50 ms and 1000 ms of audio per message, so don't send smaller chunks.
with open("webinar.pcm", "rb") as audio:
while chunk := audio.read(3200):
client.stream(chunk)
time.sleep(0.1) # pace to real time, as a live source would
client.disconnect(terminate=True)That connects to wss://streaming.assemblyai.com/v3/ws and runs on Universal-3.5 Pro Realtime, which returns partials in a few hundred milliseconds.
Three practical notes for live captions specifically.
First, don't render every partial verbatim — partials revise themselves, and a caption bar that flickers through three versions of a word is harder to read than one that lags slightly.
Second, that time.sleep is not decoration. A live microphone delivers audio in real time; a file does not. Read a file as fast as the disk allows and you dump the whole recording into the socket in milliseconds, then disconnect before the model has decided where any of the turns end. Tested against the live API, the unpaced version of this loop produced zero final turns — only partials — while the paced version produced them normally. If you're replaying a file for testing, pace it.
Third, if you need the session written out as a file afterward, buffer the final transcripts with their timestamps and serialize them to SRT or VTT at the end; the formats are trivial to generate once you have text plus a time range.
Test streaming transcription on your own audio and watch partial and final transcripts land in real time. No setup required.
Five mistakes that break subtitle files
Mixing comma and period separators. SRT uses 00:00:02,880. VTT uses 00:00:02.880. A converter that swaps the extension without rewriting the timecodes produces a file that looks right in a text editor and fails in a player.
Shipping UTF-8 with a BOM. The byte order mark shows up as a stray glyph before your first cue in several desktop players, and it can break the WEBVTT header check in a strict VTT parser because the header is no longer the first thing in the file.
Letting cues run long. A caption that stays on screen for nine seconds means the reader finished it in three and spent six wondering if the player froze. Two to seven seconds per cue is the working range, and chars_per_caption is how you enforce the upper end at export time.
Overlapping timestamps. Cue 12 ending at 00:00:14,200 while cue 13 starts at 00:00:14,000 is technically legal in SRT and undefined in practice. Some players stack both cues, some drop one. Sort and de-overlap before shipping.
Forgetting the trailing blank line. SRT blocks are separated by an empty line, and the last block needs one too. A file that ends immediately after the final text line will lose that last caption in a handful of parsers.
Accuracy is a caption problem, not just a transcript problem
It's tempting to treat format as the whole story and accuracy as someone else's department. It isn't. A caption file inherits every error in the transcript underneath it, and captions put those errors on screen where a human reads them at speed — which is exactly where they're most expensive.
Kapwing runs a browser-based video editor where creators caption their own content, and their CTO put the economics of that plainly:
“If you have an hour of content, the difference between 99% accuracy and 97% accuracy, it's a lot of time for that person to review. So you could cut down their workflow from taking half an hour, taking 20 minutes, taking 15 minutes — it's huge, right?”
— Joshua Grossberg, CTO, Kapwing
Two percentage points of word error rate sounds like a rounding error until you're the person scrubbing through an hour of timeline fixing it. The same logic is why Veed built captioning on an API rather than in-house:
“Assembly allowed our team to focus on what they are best at: Building a collaborative, browser-based video editor and distributing that product at speed and at velocity to our user base.”
— Sabba Keynejad and Tim Mamedov, Veed
If you want the deeper version of this argument, how accurate speech-to-text actually is in 2026 covers where the error budget goes and why headline WER numbers hide most of it.
One transcript, every format
The thing most people get backwards is treating the format as the output. It isn't. The transcript is the output — words, word-level timings, speaker labels, and whatever speech understanding you ran on top of it. SRT, VTT, and TXT are three serializations of the same underlying object, and once you've got the object, producing any of them is a formatting decision you can make later, or make three times.
Which means the question "SRT or VTT?" is usually the wrong question. Generate both. Store the transcript. When a platform changes what it accepts — and it will — you re-export instead of re-transcribe.
If you're evaluating which tool to generate them with rather than which format to generate, that's a different post: the best AI subtitle generators walks through the options. This one's about the files.
Transcribe once, export SRT, VTT, and plain text from the same result. Pay-as-you-go pricing, no minimums, and a free API key in under a minute.
Frequently asked questions
What is an SRT file used for?
An SRT file stores timed subtitle text for a video — a sequence number, a start and end timecode, and the caption lines, repeated for every cue. It's the most widely supported subtitle format, accepted by YouTube, Vimeo, LinkedIn, Facebook, VLC, and every major video editor, which makes it the safe default when you don't know what the destination supports.
What's the difference between SRT and VTT?
The visible difference is the timestamp separator: SRT uses a comma before milliseconds (00:00:02,880) and WebVTT uses a period (00:00:02.880). The substantive difference is capability — WebVTT adds a mandatory WEBVTT header, optional cue identifiers, metadata like language tags, CSS styling through STYLE blocks, and cue positioning. VTT is also the only one of the two that HTML5's <track> element accepts natively.
How do I create an SRT file from audio or video?
Transcribe the media with a speech-to-text API, then request the SRT export of the finished transcript rather than the JSON. With AssemblyAI that's a single call — transcript.export_subtitles_srt() in the Python SDK or a GET to the transcript's /srt endpoint — and you can set chars_per_caption to control line length. The same transcript also exports to WebVTT and plain text without re-transcribing.
Can you use an SRT file with an HTML5 video player?
No. The HTML5 <track> element only accepts WebVTT, and it fails silently when handed an SRT file — no captions appear and no error is raised. Convert the file to VTT first, which at minimum means adding a WEBVTT header line and changing the millisecond separator from a comma to a period.
What format should I use for captions on YouTube?
SRT. YouTube accepts SRT, VTT, SBV, and several broadcast formats, but it applies its own caption styling on ingest, so the positioning and CSS you'd gain from VTT get discarded anyway. Upload SRT and name the file with a language code — video.en.srt — so multi-language uploads stay organized.
Do I need timestamps if I just want the transcript text?
No, and you probably shouldn't ask for them. If the transcript is going into a summarizer, a search index, a compliance archive, or a blog post, a plain .txt export is cleaner to work with than a caption file you'd have to strip timings out of. Ask for subtitle formats when the text will play back alongside video; ask for text when it won't.