---
name: product-claims
description: The source of truth for how to describe each AssemblyAI product line in any content — web, blog, ads, email, social, decks, sales replies, docs, and FAQ. ALWAYS use this skill when writing, editing, or reviewing AssemblyAI content that names a product, model, price, language count, latency figure, or competitive claim for Pre-recorded Speech-to-Text, Realtime Speech-to-Text, Sync Speech-to-Text, Speech Understanding, LLM Gateway, Guardrails, Voice Agent API, or Dictation API. Pull the exact approved claim, and flag any copy that has drifted from it. Pairs with /medical-mode-claims for Medical Mode and /hipaa-language-check for HIPAA/BAA wording.
---

# Product claims

**Status: FIRST DRAFT — not yet reviewed by product or legal.**
Drafted 12 September 2026 from the live product pages on assemblyai.com. Every claim below is quoted or tightly paraphrased from one of these pages as they rendered that day:

| Product | Source page |
| --- | --- |
| Pre-recorded Speech-to-Text API | https://www.assemblyai.com/products/speech-to-text |
| Realtime Speech-to-Text API | https://www.assemblyai.com/products/streaming-speech-to-text |
| Sync Speech-to-Text API | https://www.assemblyai.com/products/sync-speech-to-text |
| Speech Understanding API | https://www.assemblyai.com/products/speech-understanding |
| LLM Gateway | https://www.assemblyai.com/products/llm-gateway |
| Guardrails | https://www.assemblyai.com/products/guardrails |
| Voice Agent API | https://www.assemblyai.com/products/voice-agent-api |
| Dictation API | https://www.assemblyai.com/products/dictation-api |
| All prices | https://www.assemblyai.com/pricing |

This sheet is the canonical way to describe each product so marketing, sales, and docs draw from the same numbers and language. When you write or review content, use these exact claims. If existing copy says something different, treat this sheet as the source of truth and **flag the divergence**. Don't invent a new phrasing and don't silently match the wrong one.

The Slack write-up of the contradictions, with deep links to each line, is in [`references/slack-contradictions.md`](./references/slack-contradictions.md). Where the live pages contradict each other, both versions are recorded in the **Open conflicts** register at the bottom. Until a product owner resolves a conflict, use the version marked **use this** and note the open item.

## How to apply

1. **Prices come from the pricing page**, not from FAQ copy on product pages. Product-page FAQs have drifted in at least one place (see conflict C1).
2. **Capability and positioning claims come from the product page hero and model cards.** FAQs are secondary.
3. **Precedence when pages disagree:** pricing page → product page model card or compare table → product page FAQ → meta description. Meta descriptions are the least maintained surface.
4. **Language counts are per model, not per company.** "99 languages" is a Universal-2 and platform-level claim. Universal-3.5 Pro, Universal-3.5 Pro Realtime, and Sync are 18 languages. Voice Agent API is six.
5. **Latency figures must carry their statistic.** Say "~134 ms p50", not "134 ms". Never mix a p50 and an average in one sentence.
6. **Competitive claims must be sourced.** Only repeat a competitor comparison that appears in a compare table on a live page, and link to /benchmarks where the page does. Figures on https://www.assemblyai.com/benchmarks (WER, missed entity rate, diarization cpWER, realtime latency) are quotable with named competitors when the ad or page cites that URL and dates the snapshot; the "don't cite a WER number" notes below apply to product-page paraphrase, not to benchmark-sourced copy. Paid creative follows `/ad-copywriting` for those.
7. **Brand language rules apply on top of every claim** (see the last section): Speech-to-Text not transcription, production-grade not production-ready, BAA language not "HIPAA-compliant".

---

## Platform-wide claims

These apply across product lines. Use them for company-level copy, pricing FAQ, and footers.

| Claim | Canonical version | Source and how to use |
| --- | --- | --- |
| **Company frame** | AssemblyAI builds Voice AI infrastructure: speech-to-text, streaming speech-to-text, Speech understanding, the LLM Gateway, and speaker diarization. | Brand design system intro. Audience is developers and AI/ML engineers, addressed peer-to-peer. |
| **Language breadth** | Speech-to-Text in 99 languages across our models. | Pricing FAQ says "over 99 languages across our models"; product pages say "99 languages". Use "99 languages" and attach it to the platform or Universal-2, never to Universal-3.5 Pro. |
| **Free tier** | Start free, no credit card required. The free tier includes up to 185 hours of pre-recorded Speech-to-Text and up to 333 hours of streaming Speech-to-Text. | Pricing FAQ. The standing free tier is stated in hours, not dollars. |
| **Billing model** | Pay-as-you-go, billed monthly on actual usage. No minimum commitments, upfront fees, or contracts on the pay-as-you-go plan. Speech-to-Text is billed per second. | Pricing FAQ and pricing header ("Start free, pay-as-you-go after that – no commitments required"). |
| **Multichannel billing** | Multichannel audio is billed per channel. A one-hour stereo file is billed as two hours. | Pricing FAQ. Each channel is transcribed independently, which improves multi-speaker accuracy. |
| **Procurement** | Available on the AWS Marketplace for consolidated billing. Volume discounts and custom pricing via sales. | Pricing FAQ. |
| **Enterprise tier** | Custom rate limits, enhanced concurrency, and enterprise-grade flexibility via sales. | Pricing "Custom" card on every product tab. |
| **Speaker diarization** | AssemblyAI builds and trains its speaker diarization models in-house. Available on pre-recorded and Realtime. | Pricing FAQ. Pair with Speaker Identification for real names or roles. |
| **Security** | Enterprise-grade security practices, security by design and default. | Pricing page Security block. For SOC 2, BAA, PHI redaction wording, see /medical-mode-claims and /hipaa-language-check. |
| **Customer proof (needs attribution)** | 36% improvement in close rate. 80% increase in customer satisfaction. "The new Universal-3.5 Pro speech model from AssemblyAI is best so far in terms of accuracy, latency, and language switching." | Pricing page testimonial carousel. **Customer names are not rendered in the page text.** Don't reuse these stats or the quote without confirming the attributed customer. |

---

## Pre-recorded Speech-to-Text API

**Positioning line (use verbatim):** Industry-leading accuracy on real-world audio.
**Supporting line:** Universal models top published benchmarks on noisy environments, accents, and technical vocabulary. Pick the model that fits your workload.

| Claim | Canonical version | How to use |
| --- | --- | --- |
| **Flagship model** | Universal-3.5 Pro: the most accurate, controllable model on the market. | Lead with "most accurate" plus "controllable" (natural language prompting). Pricing page adds "transcribes every conversation exactly as it's heard." |
| **Accuracy** | Universal-3.5 Pro leads our published benchmarks with industry-best accuracy on real-world audio, including noisy environments, accents, and technical vocabulary. | Always tie "industry-best" to "our published benchmarks" and link /benchmarks. Don't cite a WER number; none is on the page. |
| **Universal-3.5 Pro capabilities** | Complex, domain-specific audio. Natural language prompting. Precise entity handling. 18 languages with code-switching. Our most accurate speaker diarization yet. | Model card plus pricing card. |
| **Universal-3.5 Pro price** | $0.21/hr. | Pricing page. |
| **Value model** | Universal-2: high-accuracy Speech-to-Text at scale across 99 languages. Trained on over 12.5 million hours of audio. Exceptional accuracy at a lower price. | Pricing card. Capabilities: proven accuracy at scale, keyterms prompting (included), strong entity handling, 99 languages with code-switching. |
| **Universal-2 price** | $0.15/hr. | Pricing page. |
| **Code-switching** | Detects and transcribes code-switching in pre-recorded audio, with best results for English + Spanish or English + German. | FAQ. Don't generalize to "any language pair". |
| **Add-ons (Universal-3.5 Pro / Universal-2)** | Keyterms Prompting $0.05/hr / Included (up to 1,000 words or phrases, max six words per phrase). Prompting $0.05/hr / Not supported. Speaker Diarization $0.02/hr / $0.02/hr. Medical Mode $0.15/hr / $0.15/hr. | Pricing add-on table. Medical Mode price: see conflict C1. |
| **One-request platform** | One platform pairs the most accurate Voice AI models on the market with everything you need to turn raw audio into the outputs your product ships, on a single API request. | Features header. Features listed: Speaker Diarization, Speaker Identification, Summarization + Chapters, Sentiment Analysis, Entity Detection (50+ entity types), Topic Detection (IAB Content Taxonomy), Automatic Language Detection (99 languages), PII Redaction (text and audio in one call), Content Moderation. |
| **LLM integration** | Integrates with LLMs via LLM Gateway, a single API to OpenAI (GPT), Anthropic (Claude), Google (Gemini), and more. | FAQ. |
| **Delivery model** | Submit a job, receive results by polling or webhook. Best for long audio, batch, and latency-tolerant workloads. | Sync page compare table. Use this framing when contrasting with Sync or Realtime. |
| **Use cases named on page** | Conversation intelligence, AI notetakers, compliance monitoring, contact centers, clinical research, sales and revenue intelligence, medical transcription. | Use-case grid. |

**Don't say:** "99 languages" for Universal-3.5 Pro. "+$0.07/hr for Medical Mode" (stale FAQ, see C1). A specific WER percentage.

---

## Realtime Speech-to-Text API

**Positioning line (use verbatim):** Real-time transcription fast enough for voice agents, accurate enough for production.
**Header:** Pick the model that fits your workload.

| Claim | Canonical version | How to use |
| --- | --- | --- |
| **Flagship model** | Universal-3.5 Pro Realtime: the most accurate, controllable model on the market. The most accurate realtime model for high-quality voice agents, with built-in context carryover and conversation memory. | Model card plus pricing card. |
| **Universal-3.5 Pro Realtime capabilities** | 18 languages with code-switching. Partial and final transcripts. Highest accuracy on names, numbers, and technical terms. Self-correcting speaker labels, voice isolation, and keyterm prompting. | Model card plus pricing card. |
| **Universal-3.5 Pro Realtime latency** | ~150 ms p50 to partial and final transcripts. | Model card. **Conflicts with the FAQ's ~300 ms p50** (C4). Use the model-card figure until product confirms. |
| **Universal-3.5 Pro Realtime price** | $0.45/hr. | Pricing page. Same rate as Sync. |
| **Universal-3.5 Pro Realtime best for** | Voice agents, agent assist, notetakers, and multilingual use cases. | Model card. |
| **Value model** | Universal Streaming: production-grade streaming accuracy at the lowest price. English only. Partial and final transcripts. | Model card. Best for high-volume, English-only use cases. |
| **Universal Streaming price** | $0.15/hr. | Pricing page. |
| **Multilingual value model** | Universal Streaming Multilingual: multilingual Speech-to-Text at the speed and cost of Universal Streaming. English, Spanish, German, French, Portuguese, Italian. Language detection. | Model card plus pricing card. Best for high-volume multilingual and global contact centers. $0.15/hr. |
| **18 languages (Universal-3.5 Pro Realtime)** | English, Spanish, French, German, Italian, Portuguese, Arabic, Danish, Dutch, Hebrew, Hindi, Japanese, Mandarin, Vietnamese, Finnish, Norwegian, Swedish, Turkish. | Language grid and FAQ. Same list as Sync. |
| **Concurrency** | Unlimited concurrency: scales automatically with no hard caps on concurrent streams and no additional fees. | Compare table plus FAQ. Pricing FAQ adds the mechanics: free plan opens up to 5 new connections per minute; pay-as-you-go starts at 100 new sessions per minute and grows 10% automatically whenever you use 70% or more of the limit, with no ceiling. Custom starting limits at no extra cost. Say "no hard caps" rather than implying infinite instantaneous concurrency. |
| **Realtime features** | Immutable transcripts, intelligent endpointing, word-level timestamps, keyterms prompting (up to ~100 words), speaker diarization (10+ speakers), contextual prompting (Universal-3.5 Pro Realtime only), code-switching (Universal-3.5 Pro Realtime and Multilingual). | FAQ plus compare table. |
| **Add-ons (Universal-3.5 Pro Realtime / Universal Streaming)** | Keyterms Prompting Included / $0.04/hr. Speaker Diarization $0.12/hr / $0.12/hr. Prompting $0.05/hr / Not supported. Medical Mode $0.15/hr / $0.15/hr. Voice Focus $0.10/hr / Not supported. PII Text Redaction $0.12/hr / $0.12/hr. | Pricing add-on table. Note PII Text Redaction is $0.12/hr on Realtime but $0.08/hr on pre-recorded (Guardrails tab). |
| **Voice Focus** | Hear the speaker, not the room. Isolate the primary speaker and suppress everything else. | Pricing add-on. Universal-3.5 Pro Realtime only. |
| **Compliance** | HIPAA BAA on request. Medical Mode add-on for medical terminology. | Compare table. Use /hipaa-language-check wording. |
| **Billing** | Billed on total session duration. | FAQ. |
| **Delivery model** | Secure WebSocket. Transcripts arrive while the audio streams. Best for live captions and real-time voice agents. | FAQ plus Sync compare table. |
| **Ecosystem** | Drops into Pipecat, LiveKit, or your own stack. Available standalone or through Voice Agent API. | Voice Agents use-case card. |

**Don't say:** "~300 ms" without checking C4. "Unlimited" without the "no hard caps, scales automatically" qualifier. "Multilingual" for Universal Streaming (English only).

---

## Sync Speech-to-Text API

**Positioning line (use verbatim):** One call, one finished transcript.
**Supporting line:** Send a short clip and receive a transcript almost instantly. Universal-3.5 Pro accuracy, returned in ~134 ms.

| Claim | Canonical version | How to use |
| --- | --- | --- |
| **What it is** | A synchronous request/response Speech-to-Text endpoint. One HTTP POST in, a finished transcript back in the same response. No job to submit, no status to poll, no webhook to catch, no WebSocket to manage. | Hero plus FAQ. This "no polling, no WebSocket, no job" triad is the signature line. |
| **Model** | Powered by Universal-3.5 Pro, the same flagship model behind Pre-recorded and Realtime, not a stripped-down fast variant. | FAQ. Lead with this when accuracy is questioned. |
| **Latency** | ~134 ms p50 to a finished transcript. | Model card, pricing card, FAQ. The dictation use-case card separately cites "~160 ms average latency". Use p50 by default; if you use the average, label it as an average (C5). |
| **Limits** | Up to 2 minutes of audio per request, 40 MB maximum. Oversized requests are rejected upfront and pointed to Pre-recorded. | Model card plus FAQ. |
| **Response contents** | Per-word text, timing, and confidence in every response. | Model card. |
| **Included features** | Keyterms prompting and conversation context included at no extra cost. Prompting is $0.05/hr. | Pricing add-on table plus FAQ. |
| **Price** | $0.45/hr, the same rate as Universal-3.5 Pro Realtime. No rate limits, no upfront commitments. Volume discounts available. | FAQ. |
| **Languages** | The same 18 languages as Universal-3.5 Pro, with prompt steering. Defaults to English. | FAQ. Same list as Realtime. |
| **Not included** | No PII redaction, speaker diarization, or Speech Understanding. Use Pre-recorded or Realtime for those. | Compare-table note plus FAQ. Always state this when recommending Sync. |
| **Best for** | Dictation, voice agents with their own turn detection, short clips. Also push-to-talk, voicemail, IVR, edge functions and serverless. | Model card plus use-case grid. |
| **Dictation framing** | Speak, see it typed. A single round trip makes dictation and voice input feel instantaneous, with accuracy that on-device models can't match. | Use-case card. |
| **Three-path framing** | Every Speech-to-Text workload comes down to whether you need the answer fast, soon, or eventually. Realtime returns words live over a WebSocket. Pre-recorded submits a job for long files. Sync is the third path: one short clip in, one finished transcript back, in a single round trip. | FAQ. Use this whenever positioning the three STT APIs together. |
| **Versus Dictation API** | Sync returns a verbatim transcript, the words as spoken. Dictation API adds an LLM pass so text comes back cleaned and formatted by your instructions. Choose Sync to bring your own LLM or keep control of the raw transcript. | FAQ. |

**Don't say:** "Streaming" or "real-time" for Sync. Diarization or PII redaction as Sync features. "Instant" without a latency figure nearby.

---

## Speech Understanding API

**Positioning line (use verbatim):** Turn raw audio into structured, actionable intelligence.
**Supporting line:** Purpose-built models extract meaning from speech in a single API call.
**Frame:** The second pillar of your Voice AI pipeline.

| Claim | Canonical version | How to use |
| --- | --- | --- |
| **Model count** | Nine models, one API: Speaker Identification, Sentiment Analysis, Key Phrases, Summarization, Custom Formatting, Translation, Auto Chapters, Entity Detection, Topic Detection (IAB taxonomy). | FAQ list. **The pricing page lists eight line items** and folds chapters into Summarization ("chapter-based summaries"). See C7 before saying "nine" in pricing contexts. |
| **Single request** | Enable any combination of features in a single API call. Mix, match, and scale as your application grows. | Header plus FAQ. |
| **Speaker Identification** | Label speakers by their actual names or roles using audio context. Supports custom profiles and multi-speaker conversations. Low effort is the default; medium effort for harder recordings. | Model card plus pricing. Replaces "Speaker A / Speaker B". |
| **Sentiment Analysis** | Positive, neutral, or negative sentiment at the sentence level, and per speaker when diarization is on. | SU card plus STT features grid. |
| **Entity Detection** | Names, organizations, locations, dates, account numbers, and 50+ other entity types returned as typed fields with timestamps. | SU card plus STT features grid. |
| **Topic Detection** | Classify audio against the standard IAB Content Taxonomy. | SU card. Ideal for content moderation, routing, and analytics. |
| **Summarization** | Timestamped, chapter-based summaries of audio files at scale. Low effort default; medium effort for long, multilingual, or detail-critical audio. | Pricing card. **The SU page's Summarization card is a copy bug**: it repeats the Key Phrases description (C8). Use the pricing-page description. |
| **Auto Chapters** | Break long recordings into timestamped sections with an auto-generated summary, headline, and gist for each chapter. | SU card plus STT features grid. No separate price line (C7). |
| **Custom Formatting** | Normalize dates, phone numbers, and email addresses to machine-readable formats. | SU card. |
| **Translation** | Translate transcripts across 99+ languages in the same API request. Skip separate translation vendors. | SU card. Note "99+" here versus "99" elsewhere; use "99 languages" for consistency. |
| **Key Phrases** | Automatically extract the most significant concepts and phrases from every transcript. | SU card. |
| **Prices (per hour of audio)** | Speaker Identification $0.02 low / $0.10 medium. Translation $0.06. Custom Formatting $0.03. Entity Detection $0.08. Sentiment Analysis $0.02. Key Phrases $0.01. Topic Detection $0.15. Summarization $0.02 low / $0.07 medium. | Pricing page. You only pay for the features you enable. |
| **Audio mode** | Designed for pre-recorded audio. Pipe Realtime transcripts into Speech Understanding post-call. | FAQ. Don't position SU as real-time. |
| **Accuracy** | Built on our industry-leading Universal Speech-to-Text foundation and benchmarked regularly against real-world audio. | FAQ. No numeric accuracy claim exists for SU models. |
| **Results stats (unsourced)** | 2× conversion improvement for teams using structured conversation intelligence vs. raw transcripts. 90% reduction in manual QA review time after automating complaint and sentiment detection. <60s from audio file to fully structured intelligence in one request. | Results band. **No source or customer is cited on the page.** Don't repeat these outside the product page until a source is attached (C9). |

**Don't say:** "Real-time Speech Understanding". A numeric accuracy figure. The 2× / 90% stats in ads or sales decks without a source.

---

## LLM Gateway

**Positioning line (meta description, use verbatim):** Run open and frontier LLMs on your voice data. Fast open models for the cleanup pass, frontier models for the reasoning pass, one API key for both.
**Header:** Open and frontier models, one endpoint.
**Value header:** Easiest, most reliable way to call multiple LLMs. Ship faster, spend less on tokens, and stop losing users to provider outages.

| Claim | Canonical version | How to use |
| --- | --- | --- |
| **What it is** | A single, OpenAI-compatible API that routes requests to GPT, Claude, Gemini, Qwen, and other frontier and open models, with automatic fallbacks, no markup, and zero data retention available on every call. | FAQ. |
| **Catalog size** | 35 models from 4 providers, all OpenAI-compatible. Switch by changing one string. | Header. The four makers are Anthropic, Google, OpenAI, and open-weight (Alibaba Qwen, Google Gemma, OpenAI gpt-oss). Serving providers are AWS Bedrock, Vertex, OpenAI, and AssemblyAI. The count changes as models ship; re-check the live page before quoting. |
| **Pricing model** | 0% markup. Pay provider rates, not gateway rates. Competing gateways add 5% or more to every call. No hidden fees, no minimum commitment. | Value cards plus FAQ. **Qualifier required:** prices shown are for global routing; in-region (US or EU) routing is 10% higher due to provider cost increases. Pass `"model_region": "global"` for the listed rates. Never say "same price as the provider" without the routing qualifier (C10). |
| **Compatibility** | Drop into any OpenAI SDK. Change the base URL and API key, and the rest of your code is unchanged. | Value card plus FAQ. |
| **Fallbacks** | Configure a primary model and any number of fallbacks per request. If the primary errors, stalls, or rate-limits, the Gateway retries on the next model in your list. No code changes required. | Value card plus FAQ. |
| **Voice-native** | Your LLM calls run where your Speech-to-Text does. One less network hop on every turn. | Value card. This is the differentiator versus general-purpose gateways; lead with it in voice contexts. |
| **Model freshness** | Frontier models from OpenAI, Anthropic, and Google, plus open models we host ourselves. New models added the day they launch. | Value card plus FAQ. |
| **Self-hosted open model** | Qwen3.5 4B Fast, served by AssemblyAI in US and EU. $0.10 per 1M input tokens, $0.50 per 1M output tokens. | Catalog row. The only AssemblyAI-served model in the catalog. |
| **Demo proof point** | Dictation cleanup on Qwen3.5 4B Fast: 539 ms, 90 tokens, under $0.0001. | Hero demo. Use as an illustration of the "cleanup pass" tier, not as a benchmark. |
| **Price range** | From $0.05 per 1M input tokens (GPT-5 Nano) to $10.00 per 1M (GPT-6 Astra), global routing. | Catalog. Quote individual model prices from the live catalog only. |
| **Data residency** | EU customers can route through EU-hosted infrastructure for GDPR compliance. EU residency is on by default for traffic originating in the EU. | FAQ. |
| **Zero data retention** | Opt in per request with a header, or configure it project-wide. | FAQ. |
| **Evaluation** | Compare models before you write any code: upload a file in the playground, run a prompt over the transcript, and switch models to see how the answers differ. | Playground card. |
| **Use cases shown** | Dictation cleanup, generate content, extract info, format data. | Hero tabs. |
| **Token definition** | A token is a unit of text used by LLMs; one English word is about 1.3 tokens on average. Pricing is per input and output token of the selected model. | Pricing FAQ. |

**Don't say:** "Router" as the category (brand positioning uses "inference layer for voice AI"; see the August 2026 positioning doc). "Same price as the provider" without the global-routing qualifier. A fixed model count in evergreen copy.

---

## Guardrails

**Positioning line (use verbatim):** Profanity filtering, content moderation, and PII redaction at the text and audio level. One API, compliant by default.
**Header:** Every protection, built into your timeline. Guardrails built in, not bolted on.

| Claim | Canonical version | How to use |
| --- | --- | --- |
| **What it is** | Built-in safety features that keep voice pipelines safe and compliant: content moderation, PII redaction, profanity filtering, and operational controls like Speech Threshold, all through a single API. | FAQ. |
| **PII Text Redaction** | Identify and remove personally identifiable information, such as social security numbers and dates of birth, from the transcript before it is returned. | Feature card. $0.08/hr on pre-recorded; $0.12/hr on Realtime. |
| **PII Audio Redaction** | Replace PII spans in the original audio with a "beep" so recordings can be stored or shared for compliance reviews. | Feature card plus FAQ. $0.05/hr. |
| **Content Moderation** | Detect sensitive content in audio and video: hate speech, violence, sensitive social issues, alcohol, drugs, and more. Results include per-category confidence scores. | Feature card plus FAQ. $0.15/hr. |
| **Profanity Filtering** | Identify and redact profane or sensitive terms across any application. | Feature card. $0.01/hr. Available on all models (compare table). |
| **Operational controls** | Speech Threshold (transcribe only files with a minimum percentage of spoken audio) and Start & End of Transcript (limit output to sections containing speech) validate inputs before they hit the model. | Feature cards. No separate price listed. |
| **Implementation** | Enable via options in your Speech-to-Text request. Each feature toggles independently. No separate pipeline. | FAQ. |
| **Realtime availability** | Select Guardrails features are available on Realtime Speech-to-Text for real-time safety filtering. | FAQ. PII Text Redaction is the one priced on the Realtime tab. |
| **Language coverage** | Varies by feature. PII Redaction and Content Moderation support the core Universal model language set; see docs for per-feature availability. | FAQ. Don't claim 99 languages for Guardrails. |
| **Compliance** | Supports GDPR and PCI compliance, offers a DPA, and provides PII redaction for text and audio so regulated teams can ship safely. | FAQ. **The page's "help you meet GDPR, HIPAA, and PCI requirements" and "compliant by default" need /hipaa-language-check review** (C11). Use BAA language for anything healthcare. |
| **Competitive (compare table)** | AssemblyAI: one API, every guardrail. Deepgram: limited safety surface (no audio PII redaction, no content moderation, profanity on legacy models only). AWS: stitch your own pipeline (separate pipeline for audio PII, toxicity-only moderation, DIY profanity filter). | Compare table. Repeat only with the table's exact scope; don't extend to other competitors. |

**Don't say:** "HIPAA-compliant" or "HIPAA-ready". "Compliant by default" in healthcare copy without BAA framing. Guardrails coverage in "99 languages".

---

## Voice Agent API

**Positioning line (use verbatim):** The most accurate voice agent.
**Supporting line:** Universal-3.5 Pro gets the details right, emails, phone numbers, order IDs, and names, so your voice agent can actually complete customer tasks.
**Frame:** Your agent is only as good as what it actually hears. Invisible infrastructure for your voice product.

| Claim | Canonical version | How to use |
| --- | --- | --- |
| **What it is** | A single connection that handles speech-to-text, LLM routing, and voice generation for production voice agents. Built end-to-end on Universal-3.5 Pro Realtime. | FAQ. Pricing page adds: "A proprietary Voice AI stack, built end-to-end for production voice agents." |
| **Price** | $4.50/hr ($0.075/min), a flat hourly rate, billed per second on connected conversation time. | Pricing page. Every layer included: no per-layer add-ons, concurrency fees, or per-agent subscriptions. |
| **What's included** | Universal-3.5 Pro Realtime with prompting and Voice Focus, advanced turn detection, interruption detection, Voice Agent LLM, Voice Agent TTS, infrastructure hosting over a single WebSocket, recordings and transcripts of every conversation, SIP trunking with no markup. | Pricing FAQ. This is the full inclusion list; use it verbatim when asked "what's in the $4.50". |
| **Accuracy** | Industry-leading accuracy: lowest word error rate on real-world audio. Email addresses, phone numbers, and entity names transcribed correctly so the LLM responds to what was actually said. | Feature card. Link to /benchmarks when making the WER claim. |
| **Latency** | ~1 second end-to-end, from the user finishing a turn to the agent starting to speak. | Feature card plus FAQ. Compare table: OpenAI Realtime ~1–1.5 seconds; Deepgram not stated. |
| **Turn detection** | Speech-aware VAD. Knows when you're done talking versus pausing to think. Configurable thresholds updatable mid-conversation. | Feature card plus pricing add-on. |
| **Interruption detection** | Ignores backchannels like "mhm" and "right"; stops the agent only when the caller genuinely takes over the turn. | Pricing add-on. |
| **Voice Agent LLM** | A proprietary model tuned for spoken conversation rather than text chat. | Pricing card. Don't name an underlying model. |
| **Voice Agent TTS** | Purpose-built voices tuned for conversation, not narration. Prosody, pacing, and intonation built for real-time dialogue. Low-latency. | Feature card plus pricing card. Custom voices via sales. |
| **Session resumption** | Reconnect within 30 seconds if the WebSocket drops. Context preserved; the conversation continues where it left off. | Feature card plus FAQ. |
| **Developer experience** | Standard JSON API, no SDKs. ~6 event types (vs 30+ for OpenAI Realtime). Update prompts, voice, and tools mid-call with no reconnect. Tool calling with JSON Schema. | Feature cards plus compare table. |
| **Telephony** | Connect your own Twilio account directly. No per-minute markup on your carrier rate. Works with Twilio, LiveKit, and any telephony provider. | Pricing add-on plus use-case card. |
| **Languages** | English, Spanish, French, German, Italian, and Portuguese, with the same accuracy across all six. | Feature body plus FAQ. **The feature heading says "18 languages supported"** (C3). Use six. |
| **Competitive (compare table)** | AssemblyAI $4.50/hr on Universal-3.5 Pro. OpenAI Realtime API $18.00/hr on gpt-realtime, per-token audio billing. Deepgram Voice Agent API $4.50/hr on Nova-3, component-based billing. | Compare table. The Deepgram demo transcript on the page shows unformatted lowercase output versus AssemblyAI's formatted entities; use that contrast only with the page's own example. |
| **Enterprise** | Volume discounts, a dedicated Forward Deployed Engineer to take you live, custom voices tailored to your brand. | Pricing custom card. |
| **Clinical use** | Clinical-grade accuracy on medical terminology, with PHI redaction and a BAA available. | Use-case card. Run /hipaa-language-check. |
| **Use cases named** | Customer support, outbound sales, clinical workflows, scheduling and intake, phone agents, voice assistants. | Use-case grid. |

**Don't say:** "Production-ready" (the meta description does; it's brand-banned, C12). "18 languages" for Voice Agent API. Zero data retention for Voice Agent API (not offered; see the no-commitments campaign guardrails). "Speech-to-speech model" (it's a stack, not a single model).

---

## Dictation API

**Status: waitlist, not generally available.** All copy must preserve the early-access framing.

**Positioning line (on page):** The first API built for dictation.
**Supporting line:** Finished text back at the speed of speech. Universal-3.5 Pro accuracy, formatted by the instructions you send.

| Claim | Canonical version | How to use |
| --- | --- | --- |
| **What it is** | Speech-to-Text plus an LLM cleanup pass in a single low-latency call. Flagship accuracy with a rewrite prompt you control. | Meta description. |
| **Availability** | Join the waitlist. Early access invites go out as spots open, no commitment required. | Hero CTA plus form. Don't imply self-serve access. |
| **Speed** | Responds almost instantly. Finished text lands while your user is still looking at the screen. | Feature card. No latency figure is published; don't invent one or borrow Sync's 134 ms. |
| **Accuracy** | Built on Universal-3.5 Pro, our flagship speech model. Gets every word right. | Feature card. |
| **Formatting** | Writes the way you tell it: strips filler, applies house style, returns a structured clinical note, translates, per request. | Feature card. |
| **Languages** | Page says 19 languages. **Conflicts with Universal-3.5 Pro's 18** (C2). Use 18 until product confirms a nineteenth. | Hero. |
| **Price** | Not published. | Nothing on the product or pricing page. Don't quote a price. |
| **Build-today alternative** | Building today? Start with the Sync API: flagship Speech-to-Text in a single request, live and self-serve now. | Hero secondary CTA plus footer card. |
| **Versus Sync** | Dictation API adds the LLM pass so text comes back cleaned and formatted. Sync returns the verbatim transcript for teams that bring their own LLM. | Sync FAQ. |

**Don't say:** "The first API built for dictation" without checking the open "first API" claim conflict noted in the Dictation API launch plan (C13). Any price. Any latency number. "Available now."

---

## Quick reference

| Product | Flagship model | Price | Languages | Latency | Best for |
| --- | --- | --- | --- | --- | --- |
| Pre-recorded Speech-to-Text | Universal-3.5 Pro | $0.21/hr (Universal-2 $0.15/hr) | 18 (Universal-2: 99) | Job-based, poll or webhook | Long audio, batch, latency-tolerant |
| Realtime Speech-to-Text | Universal-3.5 Pro Realtime | $0.45/hr (Universal Streaming $0.15/hr) | 18 (Streaming: EN; Multilingual: 6) | ~150 ms p50 (see C4) | Voice agents, agent assist, live captions |
| Sync Speech-to-Text | Universal-3.5 Pro | $0.45/hr | 18 | ~134 ms p50 | Dictation, own-turn-detection agents, short clips |
| Speech Understanding | Nine models | $0.01–$0.15/hr per model | Translation across 99 | Pre-recorded only | Structured intelligence from transcripts |
| LLM Gateway | 35 models, 4 providers | Provider rates, 0% markup (global routing) | n/a | Demo: 539 ms on Qwen3.5 4B Fast | Cleanup pass + reasoning pass on voice data |
| Guardrails | Add-on features | $0.01–$0.15/hr per feature | Core Universal set | n/a | PII, moderation, profanity, input controls |
| Voice Agent API | Universal-3.5 Pro Realtime + proprietary LLM and TTS | $4.50/hr flat | 6 | ~1 s end-to-end | Production voice agents on one connection |
| Dictation API | Universal-3.5 Pro + LLM pass | Not published | 18 (page says 19, C2) | Not published | Dictation, waitlist only |

---

## Open conflicts to resolve

Each item lists what the live pages say, the version to use meanwhile, and who should confirm.

| ID | Conflict | Pages | Use this for now | Owner to confirm |
| --- | --- | --- | --- | --- |
| **C1** | Medical Mode price: pricing page says **$0.15/hr**; Pre-recorded STT FAQ says **"+$0.07/hr for Medical Mode on any base model"**. | /pricing vs /products/speech-to-text FAQ | $0.15/hr (matches /medical-mode-claims). Fix the FAQ. | Product marketing |
| **C2** | Universal-3.5 Pro language count: STT, Realtime, Sync pages say **18**; Dictation API hero says **19**. | /products/dictation-api vs the three STT pages | 18 | Dictation API PM |
| **C3** | Voice Agent API languages: feature heading says **"18 languages supported"**; body and FAQ say **six** (EN, ES, FR, DE, IT, PT). | /products/voice-agent-api | Six | Voice Agent PM |
| **C4** | Universal-3.5 Pro Realtime latency: model card says **~150 ms P50**; page FAQ says **~300 ms P50**; STT page FAQ says "within a few hundred milliseconds". | /products/streaming-speech-to-text | ~150 ms p50 (model card) | Realtime PM |
| **C5** | Sync latency: **~134 ms p50** in four places; dictation use-case card says **~160 ms average**. Not strictly contradictory (different statistics) but reads as two numbers. | /products/sync-speech-to-text | ~134 ms p50; label the average if used | Sync PM |
| **C6** | PII Text Redaction price: **$0.08/hr** on the Guardrails pricing tab; **$0.12/hr** on the Realtime add-on tab. | /pricing | Both, always stating pre-recorded vs Realtime | Pricing owner |
| **C7** | Speech Understanding model count: SU page says **nine models** including Auto Chapters; pricing page lists **eight** line items with chapters folded into Summarization. | /products/speech-understanding vs /pricing | "Nine models" on product copy; don't price Auto Chapters separately | SU PM |
| **C8** | SU page Summarization card repeats the Key Phrases description word for word. Copy bug. | /products/speech-understanding | Pricing-page Summarization description | Web team |
| **C9** | SU results band (2× conversion, 90% QA reduction, <60s) has no source or customer named on the page. | /products/speech-understanding | Don't repeat off-page | Product marketing |
| **C10** | LLM Gateway "0% markup / pay provider rates" sits beside "in-region (US/EU) pricing is 10% higher". True for global routing only. | /products/llm-gateway, /pricing | Always attach the global-routing qualifier | LLM Gateway PM |
| **C11** | Guardrails page says "compliant by default" and "help you meet GDPR, HIPAA, and PCI requirements". HIPAA wording needs legal-approved BAA language. | /products/guardrails | FAQ version: GDPR and PCI support, DPA, PII redaction; BAA language for healthcare | Legal via /hipaa-language-check |
| **C12** | "Production-ready" appears in the Voice Agent API meta description and the pricing page title. Brand-banned; should be "production-grade". | /products/voice-agent-api, /pricing | "Production-grade" | Web team |
| **C13** | Dictation API "The first API built for dictation" is an unverified category claim, already flagged in the Dictation API launch plan. | /products/dictation-api | Hold until resolved | Dictation API PM |
| **C14** | Pricing testimonials (36% close rate, 80% CSAT, Universal-3.5 Pro quote) render without customer attribution in page text. | /pricing | Confirm attribution before reuse | Product marketing |
| **C15** | Translation languages: SU page says **99+**; everywhere else says **99**. | /products/speech-understanding | 99 | SU PM |

---

## Brand language rules that override page copy

Applied on top of every claim above. Where the live page breaks one of these, the sheet's canonical version already corrects it.

- **Speech-to-Text, not transcription, for the product.** "Transcript" is fine as the output noun. Several page strings ("Realtime transcription", "transcription accuracy") are left as-is in verbatim positioning lines only; paraphrase them to Speech-to-Text in new copy.
- **Production-grade, never production-ready.** See C12.
- **HIPAA:** never "HIPAA-compliant" or "HIPAA-ready". Use "BAA available for customers processing PHI" and run /hipaa-language-check. See C11.
- **Sentence case** for headlines; product names keep their capitals (Universal-3.5 Pro, Speech Understanding, LLM Gateway, Voice Agent API).
- **Latency always carries its statistic** (p50 or average) and its scope (to partial, to finished transcript, or end-to-end).
- **Language counts are per model.** Don't let "99" attach to Universal-3.5 Pro or to Guardrails.
- **Competitive claims only from live compare tables**, linked to /benchmarks where the page links.
- **Plain, concrete copy.** State what the API does. Save cleverness for headlines.
