Extract phone call insights with LLMs in Python
Learn how to automatically extract insights from customer calls with Large Language Models (LLMs) and Python.



Call analytics is the practice of turning recorded phone conversations into structured, queryable data — transcribing the audio, attributing each turn to a speaker, and extracting fields like summaries, action items, sentiment, entities and compliance flags. This guide builds a working pipeline in Python using AssemblyAI speech-to-text for the transcript, Speech Understanding for the structured fields, and the LLM Gateway for the judgment calls a fixed schema can't make. You'll finish with JSON you can write straight into a CRM, and a clear picture of what each call costs.
What is call analytics?
Call analytics converts unstructured conversation audio into fields a system can act on. The input is a recording; the output is a row: who called, what they wanted, what was promised, how it went, and what happens next.
The argument for automating it is arithmetic, not aspiration. A contact center generating a few thousand calls a month produces hundreds of hours of audio. Human QA teams sample a small fraction of that — most reviewed calls are chosen because someone complained, which biases the sample toward the failures you already know about. An automated pipeline reads every call at a marginal cost measured in cents, which changes what questions you can ask: not "was this call bad?" but "which of the 340 calls that mentioned a competitor last month ended without a follow-up?"
That only works if the transcript is right. Every downstream classifier, phrase detector and scorecard inherits the errors in the text it was given.
Real-time vs. historical call analytics
This guide covers the historical path, which is where most conversation intelligence work lives. If you're building the batch ingestion and orchestration layer around it, the companion post on building a call center analytics pipeline in Python goes deeper on queueing, storage and dashboards than this one does.
How a call analytics pipeline works
- Ingest the recording from your telephony provider, recording store or upload endpoint.
- Transcribe it with speaker diarization so each turn is attributed to a speaker.
- Redact PII before the text leaves your processing boundary.
- Extract structured fields — entities, sentiment, topics — with purpose-built models, and bespoke fields with an LLM.
- Write the result somewhere queryable: a CRM record, a warehouse table, a dashboard.
Steps 2 through 4 are what the rest of this post builds.
Getting started
You need Python 3.9 or later and an AssemblyAI API key. The complete working code is on GitHub at AssemblyAI-Examples/extract-call-insights.
Setting up your environment
python -m venv venv
source venv/bin/activate # Windows: venv\Scripts\activate
pip install assemblyai requests
Set your API key as an environment variable rather than hardcoding it:
export ASSEMBLYAI_API_KEY=<your_key> # Windows: set
ASSEMBLYAI_API_KEY=<your_key>The sample call used throughout this post is a custom home builder inquiry: https://storage.googleapis.com/aai-web-samples/Custom-Home-Builder.mp3.
Transcribing the call with speech-to-text
Universal-3.5 Pro is the recommended default for pre-recorded audio at $0.21/hr. Pass the model ID in the speech_models list — the async parameter is plural.
import os
import assemblyai as aai
aai.settings.api_key = os.environ["ASSEMBLYAI_API_KEY"]
AUDIO_URL = "https://storage.googleapis.com/aai-web-samples/Custom-Home-Builder.mp3"
config = aai.TranscriptionConfig(
speech_models=["universal-3-5-pro"],
speaker_labels=True,
)
transcript = aai.Transcriber(config=config).transcribe(AUDIO_URL)
if transcript.status == aai.TranscriptStatus.error:
raise RuntimeError(f"Transcription failed: {transcript.error}")
print(transcript.text[:500])Two things matter here for call audio specifically.
Language coverage. Universal-3.5 Pro supports 18 languages with native code-switching, meaning it handles a speaker moving between languages mid-sentence without you declaring a language up front or splitting the audio. That is not an edge case on real customer calls — bilingual households, bilingual agents and support queues in multilingual markets switch languages inside a single utterance routinely, and a model that locks to one language at the start of the file will silently mangle the rest. If you need broader language coverage and can accept a lower accuracy tier, Universal-2 covers 99 languages total at $0.15/hr. The Universal-3.5 Pro async release notes have the full benchmark set.
Diarization quality. speaker_labels=True adds standard async diarization at +$0.02/hr and gives you transcript.utterances, each with a speaker label. Universal-3.5 Pro is optimized for cpWER — concatenated minimum-permutation word error rate — rather than DER, because what you actually care about on a call is whether the right words landed against the right speaker, not just whether the segmentation boundaries were tidy.
Format the utterances into a speaker-labeled transcript before anything downstream touches it:
def format_transcript(transcript) -> str:
return "\n".join(
f"Speaker {u.speaker}: {u.text}" for u in transcript.utterances
)
conversation = format_transcript(transcript)
If diarization is new to you, this explainer covers how it works and where it breaks.
EdgeTier, which builds AI for customer experience teams, makes the dependency explicit:
"The transcript quality is critical, both for user perception and our AI models. Once you lose trust in transcript accuracy, you erode trust in the product. For text classification, phrase detection, and agent evaluation, the language has to be correct — otherwise, the whole system falls apart."
— Dr. Shane Lynn, CEO, EdgeTier
Redacting PII before you analyze
Call recordings routinely contain card numbers, CVVs, expiration dates, home addresses, account IDs and social security numbers, usually read aloud digit by digit. If you send raw call transcripts to an LLM, a warehouse or a QA dashboard without redacting first, you have quietly expanded the scope of every system downstream.
Redaction runs as part of transcription, so the sensitive strings never exist in the transcript you store.
config = aai.TranscriptionConfig(
speech_models=["universal-3-5-pro"],
speaker_labels=True,
redact_pii=True,
redact_pii_audio=True,
redact_pii_sub=aai.PIISubstitutionPolicy.hash,
redact_pii_policies=[
aai.PIIRedactionPolicy.credit_card_number,
aai.PIIRedactionPolicy.credit_card_cvv,
aai.PIIRedactionPolicy.credit_card_expiration,
aai.PIIRedactionPolicy.banking_information,
aai.PIIRedactionPolicy.us_social_security_number,
aai.PIIRedactionPolicy.person_name,
aai.PIIRedactionPolicy.phone_number,
aai.PIIRedactionPolicy.email_address,
aai.PIIRedactionPolicy.location,
],
)
transcript = aai.Transcriber(config=config).transcribe(AUDIO_URL)
# Redacted audio is produced as a separate asset.
redacted_audio_url = transcript.get_redacted_audio_url()
Four parameters do the work:
- redact_pii turns text redaction on (+$0.08/hr).
- redact_pii_policies is the list of categories to remove. Be explicit — redacting person_name will also strip the agent's name, which you may want to keep for QA scoring.
- redact_pii_sub controls what replaces the removed span. hash writes #######; entity_name writes a readable token like [CREDIT_CARD_NUMBER], which is usually better for LLM input because it preserves the shape of the sentence.
- redact_pii_audio produces a version of the recording with the sensitive spans silenced (+$0.05/hr). Use it if the audio itself is retained for training or dispute resolution, not just the text.
Redact before the LLM step, not after. The point is to keep the raw values out of every system that isn't the transcription request itself.
Speech Understanding vs. an LLM prompt
The original version of this tutorial extracted entities and sentiment by asking an LLM to find them. That works, but it is the wrong default. AssemblyAI ships purpose-built models for exactly these tasks, and they are cheaper, deterministic and typed.
All of them are flags on the same transcription request:
config = aai.TranscriptionConfig(
speech_models=["universal-3-5-pro"],
speaker_labels=True,
redact_pii=True,
redact_pii_sub=aai.PIISubstitutionPolicy.entity_name,
redact_pii_policies=[aai.PIIRedactionPolicy.credit_card_number],
entity_detection=True,
sentiment_analysis=True,
iab_categories=True,
auto_highlights=True,
)
transcript = aai.Transcriber(config=config).transcribe(AUDIO_URL)
for entity in transcript.entities:
print(entity.entity_type, "→", entity.text)
negative = [s for s in transcript.sentiment_analysis if s.sentiment == "NEGATIVE"]
print(f"{len(negative)} negative sentences")When to use which
Use a Speech Understanding feature when the field is the same on every call. Structured features are deterministic — the same audio produces the same labels, with the same enum values, every time. They're typed, so you can put them in a database column without a validation layer. They come with character offsets and timestamps, so you can jump to the moment in the recording. They're billed per hour of audio, so cost scales with volume in a way you can forecast. And they don't drift when someone edits a prompt.
Use an LLM prompt when the output is bespoke or requires judgment. "Did the agent follow our five-step objection-handling framework?" is not an entity type. Neither is "summarize this call in the format our sales ops team uses," or "what is the single most likely reason this deal stalled?" Anything that depends on your internal rubric, changes as your business changes, or requires reasoning across the whole conversation rather than labeling a span belongs in a prompt.
In practice, do both. Redact first, extract the structured fields with Speech Understanding, then pass the redacted transcript plus the structured fields into the LLM as context. The LLM does less work, hallucinates less because the facts are handed to it, and you spend fewer tokens.
Extracting call insights with the LLM Gateway
The LLM Gateway gives you one API and one bill for models from OpenAI, Anthropic, Google and AssemblyAI's own self-hosted stack. It speaks the standard chat completions shape, so anything that talks to an OpenAI-compatible endpoint works against it.
import os
import requests
GATEWAY_URL = "https://llm-gateway.assemblyai.com/v1/chat/completions"
PROMPT = """ROLE
You are a sales operations analyst reviewing an inbound customer call.
CONTEXT
Below is a speaker-labeled, PII-redacted transcript. Structured fields already
extracted from the audio are provided alongside it. Treat those fields as ground
truth and do not contradict them.
INSTRUCTION
Produce a concise summary, the concrete action items each party committed to,
the customer's primary objection if one was raised, and a recommended next step.
If something was not discussed, say so rather than inferring it.
TRANSCRIPT
{conversation}
STRUCTURED FIELDS
{structured}
"""
def analyze(conversation: str, structured: str, model: str = "qwen3.5-4b-32k-fast"):
response = requests.post(
GATEWAY_URL,
headers={"authorization": os.environ["ASSEMBLYAI_API_KEY"]},
json={
"model": model,
"max_tokens": 1000,
"messages": [
{
"role": "user",
"content": PROMPT.format(
conversation=conversation, structured=structured
),
}
],
},
timeout=60,
)
response.raise_for_status()
return response.json()["choices"][0]["message"]["content"]The ROLE / CONTEXT / INSTRUCTION structure isn't decoration. Telling the model to treat the Speech Understanding output as ground truth measurably reduces the number of invented details, because the entity list and sentiment scores are no longer things the model has to guess at.
Choosing a model
One caveat on qwen3.5-4b-32k-fast: it mishandles forced tool_choice. If your pipeline routes results through function calls rather than parsing text or JSON, use claude-sonnet-4-6 or gemini-3.7-flash for that step. See the tool calling docs for the supported patterns.
The per-call arithmetic
Model choice looks like a rounding error on one call and like the entire budget on a hundred thousand. Work it out concretely.
A 12-minute two-party call is 0.2 hours of audio and produces a transcript of roughly 2,400 tokens. Add a ~300-token prompt and you're sending about 2,700 input tokens and getting back about 400.
Per call, with Universal-3.5 Pro, diarization, entity detection, sentiment, PII text redaction, and Qwen for the summary:
The useful way to reason about the LLM line is per million tokens against your own volume. At 10,000 calls a month you're sending roughly 27M input tokens and receiving 4M output tokens. That means every $1.00 per 1M input tokens adds about $27/month, and every $1.00 per 1M output tokens adds about $4/month. Qwen's $0.10 input rate is what keeps that line under five dollars; a model priced ten times higher moves it to roughly $32. Neither is large next to the $820 of audio processing — which is the real finding. Optimize the LLM for output quality on the calls where judgment matters, and use the cheap fast model everywhere else.
If you route requests to a specific region, add 10% for regional routing on top of the gateway line.
Getting structured JSON output for CRM integration
Asking a model nicely for JSON gets you JSON most of the time. Most of the time is not good enough for a pipeline that writes to a CRM. Use a schema contract instead: pass response_format with a json_schema and the gateway constrains generation to it.
CALL_INSIGHTS_SCHEMA = {
"type": "object",
"properties": {
"summary": {"type": "string"},
"action_items": {
"type": "array",
"items": {
"type": "object",
"properties": {
"owner": {"type": "string", "enum": ["agent", "customer"]},
"task": {"type": "string"},
"due": {"type": "string"},
},
"required": ["owner", "task"],
"additionalProperties": False,
},
},
"primary_objection": {"type": ["string", "null"]},
"call_outcome": {
"type": "string",
"enum": ["closed_won", "follow_up_scheduled", "no_next_step", "lost"],
},
"overall_sentiment": {
"type": "string",
"enum": ["positive", "neutral", "negative"],
},
},
"required": [
"summary",
"action_items",
"primary_objection",
"call_outcome",
"overall_sentiment",
],
"additionalProperties": False,
}
def extract_structured(conversation: str, structured: str) -> dict:
response = requests.post(
GATEWAY_URL,
headers={"authorization": os.environ["ASSEMBLYAI_API_KEY"]},
json={
"model": "qwen3.5-4b-32k-fast",
"max_tokens": 1000,
"messages": [
{
"role": "user",
"content": PROMPT.format(
conversation=conversation, structured=structured
),
}
],
"response_format": {
"type": "json_schema",
"json_schema": {
"name": "call_insights",
"strict": True,
"schema": CALL_INSIGHTS_SCHEMA,
},
},
},
timeout=60,
)
response.raise_for_status()
return json.loads(response.json()["choices"][0]["message"]["content"])The enum constraints are the part that pays off. call_outcome can only ever be one of four values, so the field maps cleanly to a CRM picklist and you can group by it without normalizing free text after the fact. Full details are in the structured outputs documentation.
For resilience, configure a fallback chain so a provider outage degrades to a second model instead of dropping the call from your pipeline.
Data residency, retention and compliance
Call recordings are among the most sensitive data most companies hold, so the handling model matters as much as the extraction quality.
Zero Data Retention. LLM Gateway responses are not stored. Your transcripts and the model's output don't persist on AssemblyAI's side after the request completes.
EU processing. Point the gateway at https://llm-gateway.eu.assemblyai.com/v1/chat/completions to keep LLM requests in the EU. For transcription, the EU endpoints are api.eu.assemblyai.com and streaming.eu.assemblyai.com, at the same price as US, with data staying in the EU for GDPR purposes. Regional routing carries the +10% surcharge noted above.
Certifications and BAAs. AssemblyAI holds SOC 2 Type 2, ISO 27001:2022, and PCI DSS v4.0 — the last of which is directly relevant here, since payment details spoken aloud on a support line are a routine part of call recordings. For healthcare workloads, AssemblyAI is considered a business associate under HIPAA, and offers a standard Business Associate Addendum (BAA) that is required under HIPAA to ensure that AssemblyAI appropriately safeguards protected health information. See the BAA FAQ or contact sales to get one in place.
Error handling
Three failure classes, three handlers.
import json
import assemblyai as aai
import requests
def process_call(audio_url: str) -> dict | None:
try:
transcript = aai.Transcriber(config=config).transcribe(audio_url)
if transcript.status == aai.TranscriptStatus.error:
print(f"Transcription failed: {transcript.error}")
return None
conversation = format_transcript(transcript)
return extract_structured(conversation, summarize_fields(transcript))
except aai.AssemblyAIError as e:
print(f"AssemblyAI API error: {e}")
except requests.exceptions.RequestException as e:
print(f"LLM Gateway request failed: {e}")
except json.JSONDecodeError as e:
print(f"Model returned unparseable output: {e}")
return NoneScaling call analytics in production
Polling for transcript completion is fine for one file and wrong for ten thousand. Use webhooks: submit the job, return immediately, and let AssemblyAI call you back.
webhook_config = aai.TranscriptionConfig(speech_models=["universal-3-5-pro"], speaker_labels=True, webhook_url="https://your-service.example.com/hooks/transcript-ready"); transcriber = aai.Transcriber()
future = transcriber.transcribe_async(
AUDIO_URL,
config=webhook_config,
)Your webhook handler should do one thing: validate the payload, enqueue the transcript ID, and return 200. All extraction work happens in a worker.
Store the raw structured output alongside the parsed fields. When you change a prompt or a schema, you'll want to re-run against history without re-transcribing 2,000 hours of audio.
Use cases
- Lead intelligence. Capture contact details, stated budget, timeline and next step from every inbound call, then write them to the CRM without an agent typing anything.
- Cross-call trend analysis. With consistent topic labels across your corpus, you can answer "what changed in customer objections this quarter" instead of guessing from anecdotes.
- Sales coaching. Talk ratios, objection handling and script adherence per rep, scored on every call rather than the handful a manager had time to listen to.
- Compliance monitoring. Verify required disclosures were read, and flag calls where they weren't — with content_safety catching the escalations that need human eyes.
- Support triage. Route follow-ups by sentiment trajectory and detected entities, so the calls that ended badly surface before the customer churns.
These are the workloads contact center and conversation intelligence teams build first, and they all rest on the same three-step pipeline: accurate diarized transcript, redacted and structurally enriched, then interpreted by a model.
Next steps
The pipeline in this post is short — under 150 lines of Python — because the hard parts are handled by the API. The decisions that actually shape your build are which model transcribes the audio, which fields come from structured Speech Understanding features versus an LLM prompt, and which LLM you point at the judgment layer. Get those three right and the rest is plumbing.
Frequently asked questions
How do I extract per-speaker insights from a call?
Set speaker_labels=True in your TranscriptionConfig to get diarized utterances, then format them as Speaker A: {text} before passing the transcript to an LLM or a scoring function. Universal-3.5 Pro is optimized for cpWER rather than DER, which means it's tuned for getting the right words attributed to the right speaker — the thing that actually matters when you're scoring an agent. Standard async diarization adds $0.02/hr.
How do I handle different audio formats?
The AssemblyAI Python SDK accepts most common audio and video formats without you converting anything first — MP3, MP4, WAV, FLAC, M4A, WebM and others. FLAC or 16-bit PCM WAV give the best results because they're lossless, but for typical telephony recordings the compression is already baked in upstream and re-encoding won't recover anything. Pass either a public URL or a local file path.
What's the best way to process very long phone calls?
Estimate the transcript's token count and compare it against your model's context window before you send it. qwen3.5-4b-32k-fast has a 32,768-token context, which comfortably fits most calls of an hour or less, so chunking is often unnecessary. When a call genuinely exceeds the window, split on speaker turns or fixed time intervals rather than raw character counts, summarize each chunk, then summarize the summaries — splitting mid-turn is what causes the model to lose track of who said what.
How can I get structured JSON output for CRM or database integration?
Pass response_format with a json_schema object rather than relying on prompt instructions or the looser {"type": "json_object"}. A schema lets you constrain fields to enums — so call_outcome can only ever be one of your CRM picklist values — and mark fields as required, which means the output maps to database columns without a normalization pass. The LLM Gateway constrains generation to the schema you supply, so the parse step stops being a source of pipeline failures.
How can I improve the accuracy of extracted insights?
Fix the transcript first — every classifier and extractor downstream inherits its errors, so an accuracy problem is usually a transcription problem wearing a costume. Then hand the LLM facts instead of asking it to find them: run entity detection and sentiment analysis structurally, pass those results into the prompt, and instruct the model to treat them as ground truth. Finally, structure the prompt with explicit ROLE, CONTEXT, INSTRUCTION and FORMAT sections, and tell the model to say "not discussed" rather than infer.
Should I use Speech Understanding or an LLM prompt to extract entities and sentiment?
Use a Speech Understanding feature when you need the same field on every call: entity detection, sentiment analysis, topic detection and key phrases are deterministic, typed, timestamped and billed per hour of audio, which makes them cheaper and more predictable than prompting for the same thing. Use an LLM prompt when the output is bespoke or requires judgment — a QA rubric, a summary in your team's format, a recommended next action. The strongest pipelines do both: extract structured fields first, then pass them into the LLM as ground-truth context so it reasons over facts instead of guessing at them.
How do I redact credit card numbers and PII from call recordings?
Set redact_pii=True on your TranscriptionConfig and list the categories in redact_pii_policies — credit_card_number, credit_card_cvv, banking_information, us_social_security_number and so on. Use redact_pii_sub to control the replacement (entity_name writes readable tokens like [CREDIT_CARD_NUMBER], which works better as LLM input than a hash), and add redact_pii_audio=True to silence the same spans in the recording itself. Text redaction is +$0.08/hr and audio redaction is +$0.05/hr; run both before any transcript reaches your warehouse or an LLM.
What does call analytics cost per call?
A 12-minute call transcribed with Universal-3.5 Pro plus diarization, entity detection, sentiment analysis, PII text redaction and an LLM-generated summary comes to roughly $0.083 — about $825/month at 10,000 calls. Audio processing dominates that figure; the LLM step is under $5 at the same volume using qwen3.5-4b-32k-fast at $0.10 per 1M input tokens. Add 10% to gateway costs if you use regional routing, and see pricing for current rates on every component.
Lorem ipsum dolor sit amet, consectetur adipiscing elit, sed do eiusmod tempor incididunt ut labore et dolore magna aliqua. Ut enim ad minim veniam, quis nostrud exercitation ullamco laboris nisi ut aliquip ex ea commodo consequat. Duis aute irure dolor in reprehenderit in voluptate velit esse cillum dolore eu fugiat nulla pariatur.




