Insights & Use Cases
September 30, 2026

Build a call center analytics pipeline in Python with AssemblyAI

Learn to build a Python call center analytics pipeline using AssemblyAI's Voice AI. Automatically transcribe audio, identify speakers, analyze sentiment, and create data visualizations from call recordings.

Kelsey Foster
, 
Growth
Reviewed by
No items found.
Abstract green cone illustration
Table of contents

A call center analytics pipeline turns raw call recordings into structured metrics: transcribe the audio, separate agent from customer, score sentiment per turn, and aggregate the results. This tutorial builds one in Python in about 60 lines, using AssemblyAI's Universal-3.5 Pro for transcription and speaker roles.

By the end you'll have a system that ingests a folder of recordings and emits a queryable table—one row per utterance, with speaker, sentiment, confidence, and timestamps. Everything after that is a SQL query.

What is call center analytics?

Call center analytics is the systematic analysis of customer interactions—calls, transcripts, and conversation data—to extract insight on agent performance, customer sentiment, and operational efficiency. The distinction that matters in 2026 is coverage: legacy QA programs sampled a few percent of calls because a human had to listen to each one. A transcript pipeline scores all of them.

Two words in that definition do a lot of work: systematic and interactions. Systematic rules out the spreadsheet a supervisor keeps. Interactions rules out the telephony metadata your ACD already reports—volume, wait time, transfer rate. What's left is the part nobody had access to until speech-to-text got good enough to trust, which is what was actually said.

That's the whole argument for building one. You're not replacing a manual process with a faster manual process. You're moving from a sample to a census, which changes what questions you can ask. "Which agents handle refund escalations best?" is unanswerable on a 2% sample and trivial on 100%.

The core components:

  • Speech-to-text transcription—converting audio into searchable, analyzable text.
  • Speaker diarization or channel separation—establishing who said what.
  • Sentiment analysis and entity extraction—detecting emotional tone and pulling out the things people said.
  • Aggregation and visualization—turning per-utterance rows into something a manager acts on.

If you want the strategic picture rather than the build, we've written separately about conversation intelligence and the broader set of AI use cases in contact centers.

What you'll build

By the end of this guide you'll have a working system that can:

  • Transcribe call recordings with agent and customer separated
  • Assign role labels automatically instead of generic "Speaker A" and "Speaker B"
  • Score sentiment on every conversation segment
  • Render a sentiment heatmap per speaker
  • Export structured rows for a warehouse or BI tool

System requirements

  • Python 3.9 or higher
  • A Jupyter environment (Google Colab works)
  • An AssemblyAI API key—free to start, no credit card

Install dependencies

Bash:

pip install "assemblyai>=1.0.0" pandas altair

‍

Every sample below targets Python SDK 1.0.0, released August 14, 2026, which unified the async, streaming, and sync clients behind a single client shape. If you're on 0.6x, the imports and the transcriber construction both changed—pin the version rather than guessing.

Configure authentication

Python:

import os

from assemblyai import TranscriptStatus
from assemblyai.prerecorded.v2 import Transcriber, TranscriptionConfig

# In Colab, use the secrets manager; locally, use an env var.
transcriber = Transcriber(api_key=os.environ["ASSEMBLYAI_API_KEY"])

Batch or real time?

Approach Best for Model Key features
Async (batch) Deep analysis of completed calls Universal-3.5 Pro, $0.21/hr Sentiment, entity detection, speaker roles, full-call summaries
Streaming (real time) Live agent assist during the call Universal-3.6 Pro Realtime, $0.45/hr base Partial and final transcripts in a few hundred ms, live diarization with revision (+$0.12/hr)
Sync Short IVR segments and call snippets See the pricing page One call, one response — no submit-and-poll, for clips up to a couple of minutes

This tutorial uses async. If agents need help during the call rather than after it, that's a different architecture—see real-time conversation intelligence and the streaming speech-to-text product page.

Configure the transcription request

This is the part of the tutorial worth reading slowly, because it's where most pipelines go wrong.

Python:

config = TranscriptionConfig(
    # Pin the model deliberately. Omitting speech_models means you
    # inherit whatever the platform default is on any given day.
    speech_models=["universal-3-5-pro"],

    speaker_labels=True,          # required for speaker identification

    # Sentiment is a TOP-LEVEL boolean, not a speech_understanding task.
    sentiment_analysis=True,

    # Load your product catalog, agent roster, and plan names.
    # Universal-3.5 Pro accepts up to 1,000 keyterms.
    keyterms_prompt=["Model S", "Powerwall", "Michael Johnson", "Sarah"],

    punctuate=True,
    format_text=True,

    # speech_understanding.request is a oneOf — exactly one task per request.
    speech_understanding={
        "request": {
            "speaker_identification": {
                "speaker_type": "role",
                "speakers": [{"role": "Agent"}, {"role": "Customer"}],
            }
        }
    },
)

Three things to notice.

Keyterms are doing more work than you'd think. A call center's vocabulary is a closed set—plan names, SKUs, agent first names, competitor names—and it's exactly the vocabulary a general model gets wrong. Universal-3.5 Pro takes up to 1,000 keyterms, which is enough for a whole product catalog.

Sentiment and speaker roles live in two different places, and mixing them up is a 400. sentiment_analysis is a top-level boolean. Speaker identification goes inside speech_understanding.request—and that object is a oneOf, so it takes exactly one task per request. If you need translation or custom formatting as well, that's a second request, not a second key.

The dual-channel shortcut most pipelines skip

Here's a thing worth knowing before you spend a week tuning diarization: most contact-center recordings are already stereo, with the agent on one channel and the customer on the other. If yours are, don't diarize. Channel separation is exact, and no model can beat "these are literally two different audio streams."

Python:

# Dual-channel recordings: let the channels do the separation.
config = TranscriptionConfig(
    speech_models=["universal-3-5-pro"],
    multichannel=True,
    speaker_labels=False,   # redundant when channels are already split
    punctuate=True,
    format_text=True,
)

Use diarization when you genuinely have one mixed channel—conference bridges, mobile recordings, softphone captures that downmix. Otherwise, take the free accuracy.

Transcribe the calls

A single file first, to check the shape of the output:

Python:

transcript = transcriber.transcribe("call_0001.mp3", config=config)

if transcript.status == TranscriptStatus.error:
    raise RuntimeError(f"Transcription failed: {transcript.error}")

for utt in transcript.utterances:
    print(f"[{utt.start / 1000:.1f}s] {utt.speaker}: {utt.text}")

Because speaker identification ran with speaker_type: "role", the utterances come back labeled Agent and Customer instead of A and B. One caveat worth knowing before the sentiment step: role labels are applied to utterances only. Sentiment rows still carry the raw diarization label (A, B, C), so the sentiment code below maps each row back to its role by timestamp.

One recording is a demo. A contact center is a queue, so process concurrently. The documented multi-file pattern is threads around a shared transcriber:

Python:

import threading

def transcribe_one(path, results, i):
    results[i] = transcriber.transcribe(path, config=config)

results = [None] * len(call_paths)
threads = [
    threading.Thread(target=transcribe_one, args=(p, results, i))
    for i, p in enumerate(call_paths)
]

for t in threads:
    t.start()
for t in threads:
    t.join()

Transcription is I/O-bound, so threads are the right tool and the GIL is not in your way. Serial processing on a nightly batch of a few thousand calls is the difference between a 20-minute job and an all-night one. Batch in chunks rather than spawning one thread per file if the queue is large.

Log which model actually ran

Do this. It costs one line and it will save you an incident.

Python:

# The transcript object reports the model that actually ran.
# Record it alongside every result.
print("model used:", transcript.speech_model_used)

Platform defaults change. When the async default routing shifts on September 2, 2026, anyone who didn't pin a model will see their transcripts, their diarization behavior, and their per-hour cost all move at once—with no deploy on their side. A logged model id turns that from a mystery into a one-line diff.

Build Call Analytics With AssemblyAI

Create a free account to get your API key, then run this pipeline on your own recordings. Diarization, speaker roles, and sentiment at $0.21 per hour of audio.

Sign up free

Why transcript quality is the whole pipeline

It's tempting to treat transcription as a solved commodity and spend your effort on the analytics layer. Resist that. Every classifier downstream inherits the transcript's errors, and it inherits them silently—a sentiment model given a mistranscribed negation doesn't return an error, it returns the wrong label with high confidence.

The people building on top of speech already know this. In our 2026 Voice Agent Report—455 responses, fielded Q4 2025 to Q1 2026—52.5% of respondents named accuracy and misunderstandings as their single biggest challenge, ahead of latency, cost, and integration. On the other side of the call, 55% named "having to repeat themselves" as the top end-user frustration. That survey is about live voice agents rather than post-call analytics, but the dependency is identical: whatever your system does with speech, it can only be as good as the words it was handed.

Jiminny builds a conversation intelligence platform on top of AssemblyAI, and their CEO frames the model layer as something you keep working on rather than something you buy once:

"AssemblyAI has a real high-touch personal service. It's a great partnership and we're very collaborative and get to test new AI models early and work together. And AssemblyAI is really pushing boundaries, helping us create a well-rounded conversation intelligence platform."

— Tom Lavery, CEO & Founder, Jiminny

The specific failure mode on call audio is overlapping speech. People interrupt each other constantly on the phone, and a diarization model that handles clean turn-taking falls apart on crosstalk. Universal-3.5 Pro produces the transcript and the speaker changes jointly rather than in two passes, and it's optimized for cpWER rather than DER—averaging 30.17 cpWER across meetings, telephony, far-field, and conversational audio, against Deepgram Nova-3 English at 37.92, ElevenLabs Scribe v2 at 35.26, and Gladia at 36.87.

If you want the business case rather than the metric, we've written about the true cost of inaccurate transcription.

What call center KPIs can you extract from a transcript?

This is the question a manager actually asks, and it's the part most build tutorials skip. Your phone system already knows handle time and abandon rate. What it can't see is the content of the conversation.

KPI How it is derived from the transcript Feature used
Talk-to-listen ratio Sum utterance durations per role, divide agent by customer Speaker labels and roles
Sentiment trajectory Plot per-utterance sentiment against call time; the slope matters more than the average Sentiment analysis
Dead air Gaps between the end of one utterance and the start of the next, above a threshold Word-level timestamps
Escalation language Match against a phrase list ("speak to your manager", "cancel my account") Key phrases
Script adherence Check for required disclosures in the agent's turns within the first N seconds Keyterm prompting plus role labels
PII exposure Count redaction events per call; a spike usually means a process problem, not a model one PII redaction

Sentiment trajectory is the underrated one. Average sentiment across a call tells you almost nothing—plenty of good calls start angry. The slope from first quartile to last is the number that correlates with whether the customer calls back.

Extract and structure the sentiment data

Python:

import pandas as pd

rows = []
for result in transcript.sentiment_analysis:      # SDK
    # over raw HTTP this is transcript["sentiment_analysis_results"]
    rows.append({
        "speaker": next((u.speaker for u in transcript.utterances if u.start <= result.start <= u.end), result.speaker),
        "text": result.text,
        "sentiment": result.sentiment,     # POSITIVE | NEUTRAL | NEGATIVE
        "confidence": round(result.confidence, 3),
        "start_s": result.start / 1000,
    })

df = pd.DataFrame(rows)
df["quartile"] = pd.qcut(df["start_s"], 4, labels=[1, 2, 3, 4])

That DataFrame is the actual product of this tutorial. Everything from here is charting.

Python:

import altair as alt

counts = df.groupby(["speaker", "sentiment"]).size().reset_index(name="n")

alt.Chart(counts).mark_rect().encode(
    x=alt.X("sentiment:N", title=None),
    y=alt.Y("speaker:N", title=None),
    color=alt.Color("n:Q", title="Utterances"),
).properties(width=320, height=140)

Run that across a week of calls grouped by agent and you have a QA dashboard. Run it grouped by call reason and you have a product roadmap.

See It On Your Own Call Recording

Upload a call and preview transcription, speaker roles, and sentiment in your browser. Validate the output before you wire up the notebook.

Try playground

Summarize calls without the removed parameters

If you've built a pipeline like this before, you probably reached for one of the legacy summary or chaptering parameters. Don't. They're Universal-2 only, they throw 500s against the current flagship, and they're removed on September 15, 2026. Route the summary through LLM Gateway instead, which gives you model choice and a lot more control over the output shape:

Python:

import requests

conversation = "\n".join(f"{u.speaker}: {u.text}" for u in transcript.utterances)

resp = requests.post(
    "https://llm-gateway.assemblyai.com/v1/chat/completions",
    headers={"authorization": os.environ["ASSEMBLYAI_API_KEY"]},   # no Bearer
    json={
        # Pin a current model id. The gateway publishes a retirement
        # date per model, so re-check this when you upgrade.
        "model": "claude-sonnet-5",   # or claude-sonnet-4-6, gemini-3.5-flash
        "max_tokens": 1000,
        "messages": [{
            "role": "user",
            "content": (
                "Summarize this support call in three bullets, then state "
                "the resolution status and any commitment the agent made.\n\n"
                + conversation
            ),
        }],
    },
    timeout=60,
)
print(resp.json()["choices"][0]["message"]["content"])

The gateway routes to OpenAI, Anthropic, and Google models behind one key. There are no -latest aliases—every id is pinned and carries a published retirement date—so put the model id in config and review it on a schedule. If you want to prototype the summary step for nothing, qwen3.5-4b-32k-fast is on the free tier.

Integration and deployment

Use webhooks, not polling

At a few thousand calls a day, polling is a self-inflicted wound. Pass a webhook_url on submission and let results come to you:

Python:

config = TranscriptionConfig(
    speech_models=["universal-3-5-pro"],
    speaker_labels=True,
    webhook_url="https://your-service.example.com/hooks/transcript",
    webhook_auth_header_name="X-Webhook-Secret",
    webhook_auth_header_value=os.environ["WEBHOOK_SECRET"],
)

Every major contact center platform—Five9, Genesys, Twilio—can export recordings to object storage or post them to an endpoint. Point the pipeline at that drop location and you never touch the telephony vendor's API again.

Error handling and retries

Two failure modes dominate at volume, and they need different responses. A corrupted or zero-length recording is permanent—retrying it burns quota and never succeeds, so route it to a dead-letter queue and move on. A network blip or a transient 5xx is worth retrying with exponential backoff and jitter, because the alternative is a thundering herd against the API the moment a region recovers.

The distinction matters more than the code does. Pipelines that retry everything indistinguishably are the ones that quietly stop processing at 3am and get discovered three days later when someone asks why the dashboard is flat.

Python:

import random, time

def transcribe_with_retry(path, config, attempts=4):
    for attempt in range(attempts):
        result = transcriber.transcribe(path, config=config)
        if result.status != TranscriptStatus.error:
            return result
        if "unsupported" in (result.error or "").lower():
            return None                      # permanent: dead-letter it
        time.sleep((2 ** attempt) + random.random())
    return None

What it costs

Universal-3.5 Pro is $0.21 per hour of audio for pre-recorded transcription, plus $0.02 per hour for standard speaker diarization, billed per second with no minimums. Ten thousand six-minute calls is 1,000 hours, so roughly $230 for the transcription and diarization layer.

Speech Understanding features stack on top, per hour of audio: sentiment analysis +$0.02, speaker identification +$0.02, entity detection +$0.08, PII text redaction +$0.08, custom formatting +$0.03. The pipeline in this post—transcription, diarization, speaker roles, sentiment—lands at $0.27 per hour, so about $270 for those same 10,000 calls. LLM Gateway tokens are billed separately; the full table is on the pricing page.

If you're weighing models: Universal-2 is $0.15 per hour, so the flagship costs $0.06 per hour more. On call audio that premium buys the diarization quality and the code-switching, and on a bilingual contact center it pays for itself in re-listens you don't have to do.

Worth sitting with that number for a second. Two hundred and seventy dollars to score every call your center handled in a month, against a QA program that samples one percent of them and costs a salary.

Privacy and compliance

Handle candidate and customer data deliberately. Use PII redaction to strip names, card numbers, and addresses from stored transcripts before they hit your warehouse—the redaction path covers both text and audio, so you can keep the recording without keeping the card number.

AssemblyAI enables covered entities and their business associates subject to HIPAA to use the AssemblyAI services to process protected health information (PHI). AssemblyAI is considered a business associate under HIPAA, and we offer a Business Associate Addendum (BAA) that is required under HIPAA to ensure that AssemblyAI appropriately safeguards PHI. The BAA can be signed in minutes without a sales call. If your data has to stay in the EU, api.eu.assemblyai.com is the same price as the US endpoint.

Contact center platforms including Calabrio, Concentrix, CloudCall, and AmplifAI build on AssemblyAI. Where the human review layer should sit in all this is a genuinely open design question, and the contact center solutions page covers the deployment patterns.

Where this goes next

You now have a pipeline that turns recordings into rows. The obvious extensions are trend tracking, agent benchmarking, and alerting on negative sentiment spikes.

But the more interesting move is the one that only becomes available once you have full coverage. When you scored 2% of calls, your QA program measured agents. When you score 100%, the same data measures your product—which features generate calls, which policies generate escalations, which release shipped a spike in confusion. The transcript pipeline you just built is usually sold as a contact center tool. In practice, the highest-value reader of its output is usually not in the contact center at all.

Point the first dashboard at agent performance because that's what pays for the project. Point the second one at your product backlog.

Scale This To Every Call You Handle

Talk with our team about webhooks, concurrency, redaction, and secure deployment for high-volume contact centers. We offer a BAA for covered entities.

Talk to AI expert

Frequently asked questions

What is an analytics pipeline?

An analytics pipeline is an automated sequence that moves raw data through ingestion, processing, enrichment, and output without manual steps. For call centers the raw input is audio: the pipeline transcribes each recording, separates agent from customer, extracts signals like sentiment and key phrases, and writes structured rows a dashboard can query. Once built, it runs on every new call with no human in the loop.

What are the four main types of analytics used in call centers?

Descriptive analytics reports what happened (call volume, handle time), diagnostic analytics explains why (which topics drive escalations), predictive analytics forecasts what will happen (churn risk from sentiment trend), and prescriptive analytics recommends an action (route this caller to a senior agent). A transcript pipeline is what makes the last three possible, because they all need the content of the conversation, not just its metadata.

What are the key KPIs for a call center you can get from a transcript?

Talk-to-listen ratio, dead air, sentiment trajectory across the call, first-contact resolution signals, escalation language, and script adherence can all be derived from a diarized, timestamped transcript. Traditional telephony metrics like average handle time and abandon rate come from the phone system; the transcript adds the conversational quality layer those numbers cannot see.

How accurate does transcription need to be for call center analytics?

Accurate enough that downstream classification does not inherit its errors—sentiment scoring, topic tagging, and QA scoring all degrade faster than word error rate does, because a single missed negation flips a label. Universal-3.5 Pro produces the transcript and speaker changes jointly and is optimized for cpWER, averaging 30.17 across meetings, telephony, and far-field audio, which matters for the overlapping speech common in live calls.

How much does it cost to transcribe call center audio at scale?

Universal-3.5 Pro costs $0.21 per hour of audio for pre-recorded transcription, plus $0.02 per hour for standard speaker diarization, billed per second with no minimums. Ten thousand six-minute calls come to roughly $230, or about $270 with speaker identification and sentiment analysis added. Universal-2 is $0.15 per hour if you do not need the flagship's diarization and code-switching.

How do I connect this pipeline to Five9, Genesys, or Twilio?

Every major contact center platform can export call recordings to object storage or post them to a webhook. Point the pipeline at that drop location, submit each file for async transcription with a webhook_url so results return without polling, and write the structured output back to your warehouse. No direct integration with the telephony vendor is required.

‍

Title goes here

Lorem ipsum dolor sit amet, consectetur adipiscing elit, sed do eiusmod tempor incididunt ut labore et dolore magna aliqua. Ut enim ad minim veniam, quis nostrud exercitation ullamco laboris nisi ut aliquip ex ea commodo consequat. Duis aute irure dolor in reprehenderit in voluptate velit esse cillum dolore eu fugiat nulla pariatur.

Button Text
Call Centers
Python
Tutorial