Insights & Use Cases
September 30, 2026

When keyterm prompting backfires

The reproducible cases where keyterm prompting makes transcripts worse, from overriding correct words to over-prompting, plus a pre-flight checklist.

Reviewed by
No items found.
Table of contents

Keyterm prompting is one of the highest-leverage things you can do to a transcription pipeline. Feed the model the product names, the people, the jargon that matters to your domain, and accuracy on those terms jumps. We've written about how it works for pre-recorded audio, for streaming, and in the broader prompting guide for voice agents. All of that stands.

This post is about the other direction: the cases where adding keyterms makes your transcripts worse. They're real, they're reproducible, and they tend to be discovered in the middle of a proof of concept by someone who'd been told keyterms were a free win.

Better to know the edges now.

The mental model that causes the problem

Most people's first intuition is that a keyterm list is a dictionary — a set of words the model is now allowed to output, or a lookup it consults when it's unsure.

It isn't. A keyterm list is a thumb on the scale. It shifts the probability the model assigns to the terms you supplied, across the whole transcript, including in places where the model wasn't unsure at all.

Every failure below follows from that one sentence. If you take nothing else away: keyterms bias the model everywhere, not just where it needs help.

Failure 1: overriding a word the model already got right

The most counterintuitive one, and the one that costs the most trust when a prospect finds it.

If a keyterm is acoustically close to a word that legitimately appears in your audio, the bias you added can pull the model off a correct transcription and onto your supplied term. The model heard it right. Your keyterm list talked it out of it.

Here's the cleanest demonstration, because it's the documentation's own showcase example run in reverse. The docs use a clip of someone saying "Kelly Byrne-Donoghue" that the model renders as "Kelly Byrne Donahue" without help, and show keyterms_prompt fixing it. That's the feature working exactly as advertised.

Now make the audio say Donahue for real, and leave the same keyterm in your list.

# Audio actually contains: "Hi, this is Kelly Byrne Donahue calling about the
# Halcyon account."

config = aai.TranscriptionConfig(
    speech_models=["universal-3-5-pro"],
    keyterms_prompt=["Kelly Byrne-Donoghue"],   # a name that is NOT in the audio
)

Baseline, no keyterms, three runs out of three:

Hi, this is Kelly Byrne Donahue calling about the Halcyon account.

With that one keyterm, three runs out of three:

Hi, this is Kelly Byrne-Donoghue calling about the Halcyon account.

The model transcribed the name correctly on its own, every time. Adding a similar-sounding term to the list changed it, every time — spelling, hyphenation and all. Nothing about the request said "only correct this if you're unsure."

Worth being precise about the size of the effect, because it isn't universal: the same test with a well-known word the model has strong priors for — audio saying "Vertex", keyterm "Vortex" — did not flip, three runs out of three. That's the actual shape of the risk. Keyterms win where the model was uncertain, which is exactly where a wrong keyterm also wins. Proper nouns, surnames, product names and spelling variants are the high-risk category; ordinary vocabulary mostly isn't.

Note what this means for evaluation: if you measure keyterm accuracy only on the terms in your list, this failure is invisible to you. Your metric goes up. Your transcript gets worse. You have to measure overall accuracy with and without the list, on the same audio, or you can't see it at all.

Failure 2: near-homophone collisions get worse as your list grows

Failure 1 is a special case of a general one. Every term you add is another chance to collide with something real in your audio.

Domain vocabularies are exactly the worst case here, because they're dense with similar-sounding terms by construction. Product SKUs that differ by one syllable. Competitor names that rhyme with common words. Technical jargon adjacent to ordinary English. A list assembled by exporting every term in your catalog is a list optimized to produce collisions.

The fix is scoping, not size. Ask what could plausibly be said in this specific audio, and send that:

# Weak: everything the company has ever sold
keyterms = load_all_product_names()          # ~800 terms, many near-homophones

# Better: scoped to what this conversation could contain
keyterms = (
    account.products_owned                    # what they actually use
    + account.products_in_opportunity         # what's being discussed
    + [account.name, rep.name, account.csm]   # who's on the call
)

Same mechanism, better hit rate, far fewer collisions. If you have per-call or per-file metadata — an account, a meeting title, a matter, a patient record — use it. A scoped twenty-term list will usually beat an exhaustive eight-hundred-term one.

Failure 3: it's not uniform across languages

Keyterm and context prompting are strongest in English, and the benefit does not transfer evenly to other languages. Teams running multilingual traffic have measured clear accuracy improvements in English from passing agent or conversation context, with no clear impact in some other languages in the same deployment.

The practical consequence is about expectations rather than configuration. Don't extrapolate a measured English gain across your whole multilingual footprint, and don't conclude your integration is broken when the Spanish numbers don't move the way the English ones did. Measure per language, and treat prompting as an English-first optimization that may or may not carry.

If multilingual accuracy is the actual goal, the bigger lever is the model, not the prompt. Universal-3.5 Pro has native code-switching across 18 languages built into the model rather than stitched on afterward — which addresses the underlying problem that prompting is being asked to paper over.

Failure 4: over-prompting

There's a temptation to treat the keyterm list as free — if some helps, more helps. It doesn't work that way, for three compounding reasons:

  • Collision probability rises with list length, as in Failure 2.
  • Limits are real. Universal-3.5 Pro accepts up to 1,000 words or phrases on pre-recorded audio, with a maximum of 6 words per phrase; Universal-2's older word-boost list tops out far lower; and streaming has its own, tighter limits and semantics. Know the ceiling for the model and surface you're on, and know what happens when you exceed it — silently dropping the overflow is a very different failure from rejecting the request, and it's the one that's harder to notice.
  • Signal dilutes. The fifteen terms that genuinely matter compete for influence with the seven hundred that don't.

One thing that is not a constraint, despite being widely repeated: on pre-recorded audio, prompt and keyterms_prompt are not mutually exclusive. You can send general context and a keyterm list in the same request, and both are applied — a submission with both set returns 200, completes normally, and echoes both fields back on the finished transcript. The restriction people are remembering comes from the Pipecat integration, which does enforce it. If you've built a configuration that picks one or the other on the assumption that the API rejects both, you're leaving accuracy on the table for no reason.

Test your keyterm list before you ship it

Run the same audio with and without your keyterm list in the playground and compare the full transcripts, not just the terms you supplied.

Try playground

The pre-flight checklist

Before a keyterm list goes to production:

  1. Run A/B on the same audio. Identical files, once with the list and once without. If you've never done this, do it before anything else — it's the only way to see Failure 1.
  2. Measure overall accuracy, not keyterm accuracy. Keyterm hit rate always improves. That's not the question. The question is what happened to the rest of the transcript.
  3. Scan the list for near-homophones of names and proper nouns. Say each term out loud. Surnames, product names and spelling variants are where the override risk actually lives.
  4. Scope per call or per file. Use whatever metadata you have. Twenty relevant terms beat eight hundred exhaustive ones.
  5. Check the limit for your model and surface, and what overflow does. Don't assume rejection.
  6. Measure per language. Don't extrapolate English gains.
  7. Re-run when the list changes. A list that was safe at thirty terms can collide at sixty. Treat list changes as deploys, with the same A/B gate.

The better version of this conversation

Here's the thing worth saying to a prospect who's just discovered one of these.

Every failure above comes from the same root cause: a keyterm list is a blunt, global instrument. You're telling the model "these words are more likely" with no information about when they're more likely, which is exactly why it misfires on audio where they weren't.

Context solves this at a different level. Instead of a flat list of terms, you give the model the actual situation — the question your agent just asked, the prior conversation, the domain, the document being discussed. Universal-3.6 Pro Realtime takes direction from your agent through agent_context and maintains a rolling conversation memory by default. In testing on Universal-3.5 Pro Realtime, across a benchmark of more than 10,000 voice agent audio files, passing agent context cut word error rate by 8.9%, rising to 16.4% when a context prompt is supplied alongside it, with larger reductions on the categories that matter most — place-name entities down 30.7% and fabrications down 27.0%. Contextual awareness covers the mechanism, and the Universal-3.5 Pro Realtime release post has the full numbers.

The difference is targeting. A keyterm list biases everything, everywhere. Context tells the model what's likely right now, which is both more powerful and much harder to shoot yourself with. As one partner building on it put it:

"We're excited to make AssemblyAI's Universal-3.5 Pro available on LiveKit Inference. What really stands out is their pace of innovation with Context Carryover — it intelligently applies conversation context to improve transcription accuracy in a way most speech models don't, removing the need for users to predefine key terms."

— David Zhao, Co-founder at LiveKit

Keyterms still have a place. For a fixed, small, high-value vocabulary that appears reliably in your audio, they're the simplest tool that works. Just don't treat them as free, and don't ship a list you haven't A/B tested.

Move from keyterm lists to real context

Universal-3.6 Pro Realtime takes direction from your agent and carries conversation context forward automatically. Get a free API key and try it against your own calls.

Sign up free

Frequently asked questions

Can keyterm prompting make transcription accuracy worse?

Yes. A keyterm list biases the model toward the terms you supply across the entire transcript, including in places where the model was already correct, so a term that's acoustically close to a name genuinely present in your audio can pull the model off the right answer. In testing, audio clearly saying "Kelly Byrne Donahue" came back as "Kelly Byrne-Donoghue" on three runs out of three once that variant was added to the keyterm list. This is invisible if you only measure accuracy on the keyterms themselves, because that metric improves while overall transcript quality declines. Always A/B the same audio with and without the list.

How many keyterms should I use?

Fewer than you think, and scoped rather than exhaustive. Collision probability rises with list length, and a twenty-term list built from per-call metadata — the account, the products in play, the people on the call — typically outperforms an eight-hundred-term export of your entire catalog. Universal-3.5 Pro accepts up to 1,000 words or phrases on pre-recorded audio at a maximum of 6 words each, but that's a ceiling, not a target.

Does keyterm prompting work in languages other than English?

Not uniformly. Teams running multilingual traffic have measured clear accuracy gains in English from passing keyterms or conversation context, with no clear impact in some other languages in the same deployment. Measure per language rather than extrapolating from English, and if multilingual accuracy is the real goal, native code-switching in the model is a bigger lever than prompting — Universal-3.5 Pro handles code-switching across 18 languages without configuration.

What's the difference between keyterm prompting and contextual prompting?

A keyterm list is a flat set of words biased upward everywhere in the audio, with no information about when they're likely. Contextual prompting gives the model the actual situation — the question just asked, the prior conversation, the domain or document in play — so the bias is targeted to the moment rather than applied globally. That targeting is why context misfires far less: across more than 10,000 voice agent audio files, passing agent context to Universal-3.5 Pro Realtime cut word error rate by 8.9%, and by 16.4% with a context prompt alongside it.

Can I send both a prompt and a keyterm list on the same request?

On pre-recorded audio, yes. prompt and keyterms_prompt are not mutually exclusive — a submission with both set is accepted, completes normally, and echoes both fields back on the finished transcript. The belief that they conflict comes from the Pipecat integration, which does restrict it. Check the constraint for the specific surface you're on rather than assuming it applies everywhere.

Why is my keyterm list not having any effect?

Two common causes. You may be exceeding the model's term limit and having the overflow dropped silently rather than rejected. Or the terms may already be transcribed correctly without help, in which case the list is doing nothing visible while still carrying collision risk — which is the case worth worrying about, because the risk is real even when the benefit is zero.

How do I test whether keyterm prompting is helping?

Run identical audio twice, once with the list and once without, and compare word error rate across the whole transcript rather than hit rate on the supplied terms. Do this per language and re-run whenever the list changes, since a list that was safe at thirty terms can start colliding at sixty. Treat keyterm list updates as deployments with the same testing gate you'd apply to a config change in production.

Title goes here

Lorem ipsum dolor sit amet, consectetur adipiscing elit, sed do eiusmod tempor incididunt ut labore et dolore magna aliqua. Ut enim ad minim veniam, quis nostrud exercitation ullamco laboris nisi ut aliquip ex ea commodo consequat. Duis aute irure dolor in reprehenderit in voluptate velit esse cillum dolore eu fugiat nulla pariatur.

Button Text
Universal-3.5 Pro Realtime
Word Error Rate (WER)