Insights & Use Cases
September 30, 2026

Business use cases for Generative AI

In this article, you’ll learn more about building with LLMs and the top business use cases for Generative AI tools and applications.

Reviewed by
No items found.
Abstract green mobius illustration
Table of contents

Generative AI stopped being a demo category some time ago. The question for a technical buyer in 2026 is not whether a model can write a summary — it is which workflows in your product actually justify the latency, the spend, and the error handling that come with putting a model in the path of a user.

This post walks through the business use cases that have held up in production, industry by industry, with the parameters and prices attached. Most of them share a shape: unstructured human speech goes in, structured or finished text comes out.

What is generative AI?

Generative AI is a category of artificial intelligence that creates new content — text, images, audio, code — from patterns learned in existing data. The four applications that matter commercially are text generation through language models, image synthesis, audio and music generation, and data augmentation to make downstream machine learning more robust.

For most businesses, the text branch is where the work is. It is the one that plugs into data you already have and produces output a person can act on immediately.

What are large language models (LLMs)?

Large language models are deep learning networks trained on very large text corpora and then tuned for specific applications. Anthropic’s Claude, OpenAI’s GPT, and Google’s Gemini are the best-known families. They understand complex patterns in language and generate contextually appropriate responses, which is what makes them useful for summarization, extraction, classification, and rewriting.

The practical problem is that you rarely want to bet a product on one model. Model quality moves, prices move, and providers have outages. The LLM Gateway exists for that reason: one OpenAI-compatible endpoint, 37 models as of September 2026, spanning the Anthropic, OpenAI, and Google families alongside recent additions including Kimi K3, Minimax M3, Gemma-4-31B, and GPT-6 Astra. Automatic fallback covers up to two backups with a 500 ms retry default and per-fallback overrides for prompt, temperature, and max tokens, plus prompt caching and US/EU multi-region failover.

One model in that list is worth calling out for voice work specifically. Qwen3.5 4B Fast (qwen3.5-4b-32k-fast) is the only AssemblyAI-hosted model on the roster — it runs on AssemblyAI’s own GPUs with a latency-optimized 32,768-token context, built for fast rewrite tasks like dictation cleanup, transcript rewriting, live formatting, and turn summarization. It averages 612 ms on voice rewrite tasks, 1.9× faster than GPT-4.1 and 94% cheaper per hour of audio, at $0.10 / $0.50 per million prompt / completion tokens. It is a rewrite model rather than an agentic one, so it takes max_tokens, temperature, and stream and does not support tool calling or structured-output parameters. When a rewrite sits between a user speaking and a user seeing text, that latency gap is the whole product.

Business value and ROI of generative AI

The value shows up in three places, and it is worth being precise about which one you are chasing before you build.

Productivity. Generative AI absorbs the work that is high-volume, low-judgment, and text-shaped: writing the summary, filling the field, drafting the follow-up. The gain is not that the model is smarter than the person. It is that the person was doing something that did not need them.

Customer experience. The measurable version of this is not a satisfaction survey, it is a behavior change — fewer tickets, shorter handle times, higher close rates. AssemblyAI customers publish numbers on this: Siro reports a 90% reduction in customer complaints and support tickets and a 36% improvement in close rate, and Calabrio reports an 80% increase in customer satisfaction.

Process optimization. Most organizations have a large, growing archive of recorded conversations that nobody reads. Turning that into searchable, structured, queryable data is the least glamorous generative AI use case and frequently the highest-return one.

A note on cost modeling, because it is where projects stall. Token-metered pricing makes forecasting hard: the bill depends on how much people talk and how verbose the model is, neither of which you control precisely. Flat per-hour-of-audio pricing removes that variable. It is why the voice products below are priced per hour with features included rather than per token.

Industry-specific generative AI use cases

Healthcare: transform patient care and clinical workflows

Clinical documentation is the canonical case. A clinician talks — to a patient, or to a recorder after the visit — and a model produces the note, the orders, and the follow-up message. The value is measured in minutes per encounter returned to the clinician, multiplied by encounters per day.

Accuracy on this workload is not general accuracy. It is accuracy on drug names, dosages, anatomy, and procedure terms — the tokens where an error is clinically meaningful. Medical Mode tunes the model for exactly that vocabulary and is activated with a single parameter, domain: “medical-v1”, paired with speech_models: [“universal-3-5-pro”] on the async path or speech_model: “universal-3-5-pro” on streaming. It posts a 3.2% Missed Entity Rate and 87% fewer entity errors than the base model. It covers English, Spanish, German, and French, and costs $0.15/hr on top of the base model.

Contextual prompting is the complementary lever, and it stacks rather than competing. Medical Mode teaches the model clinical vocabulary in general; contextual prompting tells it about this encounter. On a public benchmark across 20,000 real voice-agent calls, detailed context cut medical-term entity errors by 43% and scenario-level context by 24%. Internal testing on prior-visit notes fed as context showed a 31% reduction in missed medical terms.

On compliance, the accurate framing is this: AssemblyAI enables covered entities and their business associates subject to HIPAA to use the AssemblyAI services to process protected health information (PHI). AssemblyAI is considered a business associate under HIPAA, and we offer a standard Business Associate Addendum (BAA) that is required under HIPAA to ensure that AssemblyAI appropriately safeguards PHI. The BAA can be reviewed and signed self-serve without a sales call. PHI redaction runs across audio and transcripts, and the platform carries SOC 2 Type 2, ISO 27001:2022, and PCI DSS v4.0.

Companies building here include Sully AI, Heidi Health, DeepScribe, Knowtex, and Commure.

“We’ve integrated the newest models from AssemblyAI for pre-recorded audio ASR in our ambient product, and it’s been excellent…”

— Gautam Pradeep, Tech Lead, Commure

Test Medical Mode On Your Own Audio

Paste in a dictated chart note and compare output with and without Medical Mode before you write any integration code.

Try playground

Customer service: analyze interactions and improve experiences

Contact centers generate enormous volumes of unstructured audio that historically got sampled at single-digit percentages by QA teams. Generative AI changes the economics of reading all of it: transcribe every call, run sentiment analysis and topic detection across the corpus, then generate per-call summaries and per-agent coaching in the format the business already uses.

The pattern that works is two-stage. Speech-to-text produces the transcript and the Speech Understanding features produce structured signals; an LLM pass then turns those into the artifact a human wants — a summary, a disposition code, a coaching note. Splitting it this way means the expensive model only sees what it needs to.

Conversation intelligence and contact center platforms building on this include CallRail, WhatConverts, Jiminny, EdgeTier, Calabrio, Concentrix, and CloudCall.

Financial services: automate compliance and risk analysis

Financial institutions record conversations because they are required to, and then largely leave them unexamined. Generative AI turns that obligation into an asset: automated compliance monitoring against policy language, audit trails generated from the recording rather than from memory, sentiment-based risk flags, and real-time transcription on trading floors.

The output is the point — stored conversations become structured, auditable intelligence that a compliance team can query rather than sample. PII redaction matters here for the same reason PHI redaction matters in healthcare: the transcript is a new copy of sensitive data and has to be treated as one.

Content creation: accelerate video and marketing workflows

Content teams are producing more surfaces than they can staff, and most of what they produce starts as a recording — a webinar, a podcast, an interview, a product walkthrough. Transcription plus an LLM pass turns one recording into the transcript, the summary, the SEO description, the chapter markers, the social clips, and the highlight reel.

The economics work because the marginal cost of the seventh derivative asset is a model call, not a person. Teams building in this space include Veed.io, Descript, HeyGen, Happy Scribe, and Runway.

Education: enhance learning management systems

Learning platforms use transcription first for accessibility — captions and searchable transcripts for recorded lectures — and then extend it into generative features: post-session summaries, automatically generated study guides, and question-answering over course material. The accessibility requirement pays for the infrastructure; the generative features are what students actually notice.

Legal: streamline case preparation and documentation

Legal teams transcribe depositions and hearings automatically, apply sentiment analysis to surface moments worth a second look, and replace linear manual review with targeted querying over the transcript. Speaker attribution matters more here than almost anywhere else, because a quotation attributed to the wrong person is worse than no quotation. Averaged across DiPCo, CALLHOME, NOTSOFAR and AMI, Universal-3.5 Pro posts cpWER 30.17, with Azure closest at 30.35, then ElevenLabs Scribe V2 at 35.26, Speechmatics at 36.60, Gladia at 36.88 and Deepgram at 37.93.

Dictation: turn spoken input into finished text

Dictation is the cleanest illustration of what generative AI adds to speech, because the gap it closes is so visible. Transcription models are verbatim by design — but what was said is not what anyone wants to send. Real speech contains filler, restarts, and self-corrections: “um so can we uh move the the meeting to thursday i think friday works better actually.” A perfect transcript of that is still not a message.

The Dictation API puts speech-to-text and an LLM rewrite in a single call and returns “Can we move the meeting to Friday? That works better.” — filler gone, the self-correction resolved to what the speaker landed on, tone intact. The verbatim transcript comes back alongside it in the text field, so nothing is lost. An llm_instruction parameter lets you replace the default cleanup with your own plain-English rewrite task, which is how the same endpoint produces a chart note for a clinician and a Jira ticket for an engineer.

The business case is the billing model as much as the capability. It is $0.62/hr flat, and per the pricing page: “Every feature is included in the $0.62/hr rate. One line on your bill covers the whole request — there is no second rate to model and no token math to do.” Building the same thing yourself means an STT bill, an LLM bill metered in tokens you cannot forecast, and the orchestration between them. It returns finished text in about 0.36 s. The constraints to design around are a 120-second cap per call and WAV or raw PCM input only. Details are in the Dictation API docs.

Voice agents: build real-time conversational AI

Voice agents are the most demanding generative AI deployment most companies will attempt, because every component runs in the latency budget of a human conversation. The applications are well established: customer support automation, real-time assistance for human agents, appointment scheduling, and IVR modernization. Gartner predicts that by 2029, agentic AI will autonomously resolve 80% of common customer service issues.

What has changed is the integration cost. The traditional build is three vendors — speech-to-text, LLM, text-to-speech — three bills, three failure modes, and a latency budget you assemble by hand. The Voice Agent API collapses that into one WebSocket at wss://agents.assemblyai.com/v1/ws, bundling Universal-3.6 Pro Realtime, a Voice Agent LLM, and Voice Agent TTS behind a single connection with a small minimum event loop. End-to-end latency is about one second.

Pricing is flat at $4.50/hr ($0.075/min). From the pricing page, verbatim: “Every feature is included in the $4.50/hr rate. There are no per-layer add-ons, concurrency fees, or per-agent subscriptions.” Billing is per second, with no minimums and no commitment — which is the part that matters when your call volume is spiky.

On accuracy under real conditions, AssemblyAI’s English voice-agent benchmark runs 12,460 scripted voice-agent scenarios rather than read speech. Universal-3.6 Pro Realtime posts a 5.19% word error rate against Deepgram Flux EN at 13.50%, ElevenLabs Scribe v2 at 7.78%, and Deepgram Nova-3 at 8.64%. On entity error rate — the names, codes, and phone numbers an agent has to get exactly right — it posts 14.4% against Flux EN at 30.1%, Scribe v2 at 18.5%, and Nova-3 at 26.1%.

Two features are doing unusual work there. agent_context lets you pass the agent’s own spoken reply into the stream so that a short or mumbled response resolves against what was said; across a benchmark of 10,000+ voice agent audio files it cut WER by 8.9%, rising to 16.4% with a context prompt on top, with fabrications down 27.0% and place-name entity errors down 30.7% on that combined configuration. voice_focus tunes for the acoustic environment, with near-field for headsets and phones and far-field for rooms, kiosks, and drive-thrus. Turn detection does not run on a silence timer — when a speaker pauses, the model evaluates what has been said so far to judge whether the turn is complete, which is what keeps it from splitting a phone number across two turns.

Companies building voice agents on this stack include Decagon, Bland AI, Phonely, Vapi, Synthflow, Lorikeet, and Zowie.

Scoping A Production Voice Agent

Talk through latency budgets, flat per-hour pricing, accuracy on your own call audio and deployment options with someone who has shipped these before.

Talk to AI expert

Customer success stories and implementation results

The clearest signal about which use cases are real is which ones companies have shipped and published numbers on.

In conversation intelligence, Siro reports a 90% reduction in customer complaints and support tickets and a 36% improvement in close rate, and Calabrio reports an 80% increase in customer satisfaction. Both numbers are published on the pricing page.

In content creation, Veed.io, Descript, and HeyGen build transcription and generative features directly into their editing products. In meetings and collaboration, Granola, Supernormal, Fireflies, Metaview, and Dovetail turn recorded conversations into structured notes. In healthcare, Sully AI, Heidi Health, DeepScribe, and Commure build clinical documentation products on the same infrastructure. Zoom uses AssemblyAI in AI R&D.

Getting started with generative AI implementation

The framework that survives contact with a real roadmap is unglamorous.

Start from a problem with a number attached. “Add AI” is not a project. “Cut average note-writing time from nine minutes to three” is, and it tells you when to stop.

Build the proof of concept against your own audio. Benchmark numbers are a filter, not an answer. Run your actual recordings — the noisy ones, the accented ones, the ones with three people talking — through the playground before you write integration code. The free tier covers 185 hours of pre-recorded and 333 hours of streaming transcription, which is more than enough to settle the question.

Measure before you scale. Instrument the metric you named in step one and compare against the pre-AI baseline. If you did not capture the baseline, capture it now.

Choose infrastructure on the terms that bite later. Accuracy on your domain, latency under your conditions, documentation you can actually follow, pricing you can forecast, deployment options that match your compliance posture — hosted, self-hosted in your own VPC, or EU data residency at api.eu.assemblyai.com with no price difference.

Ready To Build?

Get 185 hours of pre-recorded and 333 hours of streaming transcription on the free tier, and run your own audio through the same infrastructure these use cases are built on.

Sign up free

Frequently asked questions about generative AI business use cases

What is the best speech-to-text API for building voice agents?

A voice agent needs three things from speech-to-text: very low latency, high accuracy on names and domain-specific terms, and turn detection that knows when the user has actually finished. On AssemblyAI’s English voice-agent benchmark of 12,460 scripted voice-agent scenarios, Universal-3.6 Pro Realtime posts a 5.19% word error rate and a 14.4% entity error rate — ahead of Deepgram Flux EN, Deepgram Nova-3, and ElevenLabs Scribe v2 on both. Turn detection decides from what has been said rather than from a silence timer, which is what keeps it from cutting a caller off mid-sequence.

How can businesses measure the ROI of generative AI?

Track metrics tied to the business goal rather than to the model: reduced operational cost, increased throughput per person, higher customer satisfaction, and new revenue from AI-powered features. Capture the pre-deployment baseline before you launch, because reconstructing it afterward is guesswork. Published examples include Siro’s 90% reduction in customer complaints and support tickets and 36% improvement in close rate, and Calabrio’s 80% increase in customer satisfaction.

What speech-to-text API is recommended for healthcare applications?

Healthcare needs accuracy on medical terminology, reliable speaker separation for provider–patient conversations, and a vendor that will sign a Business Associate Addendum. AssemblyAI is a business associate under HIPAA and offers a standard BAA, which can be reviewed and signed self-serve without a sales call. Medical Mode activates with domain: “medical-v1” and posts a 3.2% Missed Entity Rate — 87% fewer entity errors than the base model — for an additional $0.15/hr.

What is AssemblyAI’s LLM Gateway and how does it work?

The LLM Gateway is a single OpenAI-compatible endpoint that gives you 37 models as of September 2026, spanning the Anthropic, OpenAI, and Google families plus recent additions including Kimi K3, Minimax M3, Gemma-4-31B, and GPT-6 Astra. It handles automatic fallback across up to two backup models with a 500 ms retry default, prompt caching, and US/EU multi-region failover. For latency-sensitive voice work it also offers Qwen3.5 4B Fast (qwen3.5-4b-32k-fast), the only AssemblyAI-hosted model on the roster, averaging 612 ms on voice rewrite tasks — 1.9× faster than GPT-4.1 and 94% cheaper per hour of audio.

How is generative AI different from other types of AI?

Traditional AI recognizes patterns and makes predictions — classifying a call as at-risk, or detecting an entity in a transcript. Generative AI produces new content: text, summaries, images, or code. Most production systems use both, with the discriminative model producing structured signals and the generative model turning them into something a person reads.

Which industries benefit most from generative AI implementation?

Industries with high volumes of unstructured data see the greatest returns — healthcare, financial services, customer operations, content creation, education, and legal. The common factor is a large archive of recorded human conversation that was previously too expensive to read, which generative AI makes economical to process end to end.

Title goes here

Lorem ipsum dolor sit amet, consectetur adipiscing elit, sed do eiusmod tempor incididunt ut labore et dolore magna aliqua. Ut enim ad minim veniam, quis nostrud exercitation ullamco laboris nisi ut aliquip ex ea commodo consequat. Duis aute irure dolor in reprehenderit in voluptate velit esse cillum dolore eu fugiat nulla pariatur.

Button Text
Product Management
Generative AI