Insights & Use Cases
September 30, 2026

Building with transcripts: Search, indexing, display and downstream Integrations

Transcript search explained: learn how indexing, timestamps, display, and integrations make audio and video searchable, accurate, and easy to use at scale.

Kelsey Foster
, 
Growth
Reviewed by
No items found.
Abstract green diamond illustration
Table of contents

Nobody searches a transcript archive for the word "the."

They search for a company name, a person, a product, a dollar figure, a city, a date. Almost every real query against a transcript library is a proper noun or a number — which means the thing that determines whether your search works isn't your average word error rate. It's whether the model got the entities right.

That distinction changes what you build and what you measure. Here's the version of transcript search that holds up in 2026.

How does transcript search work?

Four layers: transcribe the audio with word-level timing, extract structure from the transcript, index both, then query and render results that link back to the moment in the audio.

The temporal metadata separates this from ordinary document search. A hit isn't useful as a page of text — it's useful as a timestamp someone can click to hear the sentence in context, so every layer downstream has to carry start and end through unchanged.

Where teams go wrong is treating the transcript as a flat blob of text to be tokenized. The transcript is a structured object: words with timings and confidence, speaker turns, and — if you ask for it — entities, topics, and key phrases already extracted. Indexing only the string throws most of that away.

Universal-3.5 Pro returns all of it in one request, which matters more than it sounds: every extra pass you run over the audio is another place for the index to drift out of sync with the source.

What accuracy actually matters for search?

Entity and proper-noun accuracy. Not aggregate word error rate.

Think about the failure mode. A transcript that renders "quarterly revenue" as "courtly avenue" scores as two errors out of however many thousand words — statistically trivial, functionally fatal, because that segment is now unfindable by the only query anyone would use to find it. Meanwhile, a model that mangles filler words and gets every company name right will feel near-perfect to a user searching a call archive.

Aggregate WER averages those two cases together and tells you nothing. Entity error rate tells you what you need to know. Here's the clearest published comparison, from AssemblyAI's English voice-agent benchmark (12,460 scripted voice-agent scenarios) with the realtime model:

Metric (lower is better) Universal-3.6 Pro Realtime Deepgram Flux EN ElevenLabs Scribe v2 Deepgram Nova-3
Entity error rate 14.4% 30.1% 18.5% 26.1%
Names 10.9% 29.0% 14.8% 24.3%
Codes / IDs 10.0% 46.1% 12.0% 27.6%
Phone numbers 2.4% 11.5% 3.4% 4.5%

Source: Universal-3.6 Pro Realtime research.

The gap between 14.4% and 30.1% is the gap between a search box people use and one they abandon.

You can push entity accuracy further on your own vocabulary. Keyterm prompting takes a list of the names, products, and jargon specific to your corpus and biases the model toward them. On async you can supply up to 1,000 words or phrases, at a maximum of six words per phrase — treat that as a ceiling rather than a promise, since actual capacity may be lower due to internal tokenization, each word in a multi-word phrase counts toward the limit, and capitalization and longer words consume more capacity. It's a +$0.05/hr add-on on async, and it's included at no extra cost on streaming speech-to-text with Universal-3.6 Pro Realtime — where the ceiling is different again: 100 keyterms per session, each string 50 characters or shorter.

Keyterms aren't your only lever. The free-text prompt describes the situation — what kind of recording this is, who's on it, what it's about — while keyterms_prompt enumerates the exact strings you want the model biased toward. For a search corpus the term list usually does the heavy lifting, since your queries are proper nouns, but a sentence of context costs nothing to add.

Test entity accuracy on your own corpus

Run a handful of your hardest recordings — the ones full of client names and product codes — and check whether the terms you'd actually search for come back intact.

Sign up free

What should you index besides the words?

Entities, topics, and key phrases — extracted at transcription time, not derived later with regex.

This is the biggest change to how transcript search gets built. Speech understanding returns structured fields alongside the transcript in the same request:

  • Entity detection across 50+ types — people, organizations, locations, dates, money, phone numbers, email addresses, medical terms. These become filterable facets rather than strings you hope match.
  • Topic detection against the IAB taxonomy, which gives you category browsing without training a classifier.
  • Key phrases (English), useful as a lightweight relevance signal and for result snippets.
  • PII redaction, if the index shouldn't contain what the audio contained.

Indexing on typed entities rather than raw tokens is what lets a user ask for "every mention of Acme Corp by a customer, last quarter" instead of grepping for a string that might be a company, a person's surname, or a typo.

Index structure and schema design

Segment-based indexing beats document-based for anything conversational. Chunking on speaker turns works better than fixed 30-second windows, because a turn is a semantically coherent unit and a 30-second window is an arbitrary one that regularly slices a sentence in half.

A workable segment document:

  • text — the segment's spoken content
  • start / end — millisecond timings, carried through untouched
  • speaker — from diarization
  • entities[] — typed, with their own offsets
  • topics[] — IAB labels for the parent transcript
  • confidence — for down-ranking low-certainty segments
  • file_id, recorded_at, and whatever your business metadata is

Real-time vs batch indexing

Batch for archives, real-time for live products. If users need to search a call while it's still happening, index streaming turns as they arrive and accept that the last few segments may be revised. Otherwise, index on the completion webhook and get better compression for it.

See the structured output before you design your schema

Upload a recording and look at the raw JSON — entities, topics, speaker turns, word timings — so your index maps to what the API actually returns.

Try playground

How do you display transcript search results?

Show enough context that the hit is interpretable, and make the timestamp clickable. Everything else is detail.

Speaker-turn context windows beat fixed sentence counts or fixed durations — show the full turn containing the hit, plus the turn before and after, so a "yes" doesn't appear stranded without the question it answered.

Word-level timestamps let you do the thing that makes people trust the search: highlight the matched word inside the segment, and start playback a second or two before it, so the listener hears the run-up rather than dropping into the middle of a syllable. That small detail does more for perceived quality than most ranking work.

Can you do semantic search over transcripts?

Yes, and in 2026 you probably should — keyword indexing alone can't answer "which calls had pricing objections" when nobody said the word "objection."

The practical setup is a hybrid. Keep the keyword and entity index for precise lookups, where a user knows the exact name or number they want. Add a vector index over the same speaker-turn segments for conceptual queries. Run both, merge the rankings.

For the generative layer on top — question answering across a library, summarized answers with citations back to timestamps — LLM Gateway gives you one API into OpenAI, Anthropic, Google, and others. The retrieval pattern is standard RAG, with one transcript-specific advantage: your chunks already carry speaker labels and timings, so every generated answer cites the exact moment it came from. That's a much stronger citation than a page number, and it's the foundation most conversation intelligence products are built on.

Financial research platform AlphaSense runs search across a transcript library at genuine scale; meeting infrastructure provider Recall.ai supplies the recordings that many transcript search products index in the first place. Same architecture in both cases — capture, structure, index, retrieve.

How much storage does a transcript index need?

Less than the audio by orders of magnitude, and more than you'd guess from the text alone.

Raw transcript text is tiny. What grows is everything around it: word-level timings roughly double the payload, entities and topics add more, and a vector index is typically the largest single component once you're storing embeddings per speaker turn. Size the keyword index and the embeddings separately — they scale on different things, and the embeddings are the ones that will surprise you.

The thing worth internalizing

Transcript search fails on nouns, not on averages.

Every optimization that matters follows from that — measuring entity error rate instead of WER, priming the model with your own vocabulary, indexing typed entities rather than tokens, chunking on speaker turns, keeping timestamps intact all the way to the play button. Get those right and the search box feels like it's reading the user's mind. Get them wrong and no amount of ranking cleverness saves you, because the word they're looking for isn't in the index at all.

Request shapes and the full field reference are in the API reference; per-second rates for transcription, diarization, and speech understanding add-ons are on pricing.

Build your transcript index on Universal-3.5 Pro

Entity accuracy that holds up on names and numbers, structured output in the same request, and per-second billing with no minimums.

Sign up free

Frequently asked questions

How do I perform search on transcribed audio content?

Transcribe with word-level timestamps and speaker labels, extract entities and topics in the same request, then index each speaker turn as a document carrying its text, timings, speaker, and typed entities. Query that index and render each hit as a snippet with a clickable timestamp that seeks the audio. Indexing typed entities rather than raw tokens is what makes precise lookups work, since most real queries are proper nouns.

What is the difference between semantic and keyword search over transcripts?

Keyword search matches the literal strings a speaker used, which is what you want when a user knows the exact name, product, or number they're looking for. Semantic search matches meaning through embeddings, which is what you need for conceptual queries where the relevant words were never spoken, such as finding calls containing pricing objections. Most production transcript search runs both over the same segments and merges the rankings.

How to transcribe segments of a video with API calls?

Submit the full file and use the word-level timestamps in the response to slice out the segments you need, rather than pre-cutting the media. This keeps context intact for the model, which improves accuracy at segment boundaries, and it gives you timings that line up with the original file for playback. If you only need a short clip transcribed immediately, the Sync API returns results in a single request for audio up to two minutes.

Does transcript search work across multiple languages?

Yes, though the model choice depends on coverage. Universal-3.5 Pro handles 18 languages with native code-switching, which matters for transcript search because multilingual speakers switch mid-sentence and a search index built on a monolingual pass will lose those segments. For coverage beyond those 18, Universal-2 supports 99+ languages and remains fully supported.

How much storage does a transcript search index need?

Far less than the source audio, but plan for the metadata rather than the text. Word-level timings, entity annotations, and speaker labels together often outweigh the transcript string itself, and a vector index for semantic search is usually the largest component once embeddings are stored per speaker turn. Since default transcript retention drops to 30 days with deletions beginning September 30, 2026, budget for holding the full transcript objects yourself.

How to get started using a speech-to-text API for audio transcription?

Create an account, get an API key, and submit a file with speaker_labels enabled and a keyterm list covering the names and jargon in your corpus. Omitting the model field auto-selects the latest Universal Pro model, so most integrations do not need to pin a version. Results arrive via polling or a webhook, and the same request can return entities, topics, and key phrases alongside the transcript.

Title goes here

Lorem ipsum dolor sit amet, consectetur adipiscing elit, sed do eiusmod tempor incididunt ut labore et dolore magna aliqua. Ut enim ad minim veniam, quis nostrud exercitation ullamco laboris nisi ut aliquip ex ea commodo consequat. Duis aute irure dolor in reprehenderit in voluptate velit esse cillum dolore eu fugiat nulla pariatur.

Button Text
Speech-to-Text