Insights & Use Cases
August 25, 2026

How to build an AI-Powered interview scoring system with speech-to-text

This tutorial shows you how to build an AI-powered interview scoring system that records interviews, converts speech-to-text, and systematically evaluates candidates using structured criteria.

Kelsey Foster
, 
Growth
Reviewed by
No items found.
Abstract green cone illustration
Table of contents

An interview scoring system is a structured tool hiring teams use to rate candidates against fixed criteria instead of gut feel. An AI-powered one adds the evidence layer: it transcribes the interview, then attaches the candidate's actual words and timestamps to every score, so a rating can be audited rather than trusted.

That second sentence is the whole point of this post. Plenty of tools will give you a 1–5 scale and a template. Almost none of them can answer the question that actually matters six weeks later, when a rejected candidate's manager asks why they scored a 2 on system design.

This tutorial builds the version that can. You'll use Python and AssemblyAI's speech-to-text API to transcribe interviews with speaker separation, then extract competency evidence with timestamps and generate a scorecard where every rating links to a quote.

What is an AI-powered interview scoring system?

It's a structured evaluation tool that records the interview, converts the audio to text, and then searches that text for evidence of specific skills. The practical effect is that you stop doing two jobs at once.

Traditional scoring happens while you're talking. You're listening, taking notes, and evaluating simultaneously, and your brain is bad at all three under load. AI-powered scoring separates them completely: interview first, score later, against the complete transcript.

Manual scoring vs transcript-based scoring

Dimension Manual scorecard Transcript-based scoring
Evidence behind a score A paraphrase written from memory A verbatim quote with a timestamp you can replay
Consistency across a panel Each interviewer scores from their own notes Everyone reviews identical source material
Auditability of a disputed decision Reconstructed after the fact The record already exists and is timestamped
Interviewer attention during the call Split between listening, writing, and judging Entirely on the conversation
Recency bias High — the last ten minutes dominate Low — the whole interview is equally accessible
Recalibrating past decisions Impossible Re-run the new rubric over the existing transcripts

That last row is the one people underestimate. If you change your rubric in Q3, a transcript archive lets you re-score Q1's candidates against it. A folder of handwritten notes doesn't.

The four components

Every scoring system needs four parts working together:

  • Scoring criteria—four to six job-specific skills you can actually observe in an answer.
  • A rating scale—1 to 5, with a written behavioral anchor for each number.
  • Evidence extraction—direct quotes that prove or disprove each competency.
  • Complete transcripts—the single source of truth, with speaker labels and timestamps.

The fourth one is what unlocks the other three. Once every interview is searchable text, you can grep a panel's worth of conversations for a phrase, hand a transcript to a second reviewer for an independent read, and build the same kind of searchable evidence layer teams already run over sales calls in conversation intelligence tools.

Get Speaker-Labeled Interview Transcripts

Create a free account, grab an API key, and run the Python below on a real interview recording. Diarization and timestamps make evidence quotes effortless.

Sign up free

How to build an AI-powered interview scoring system

Five steps. Each builds on the previous one.

Step 1: Define role-specific scoring criteria

Pick four to six skills that actually predict success in the specific role. Not "good communication"—that means completely different things for a backend engineer and an enterprise AE.

For a software engineering role, look for observable behaviors: how the candidate decomposes an ambiguous problem, how they reason about system design trade-offs, depth in the relevant stack, and whether they can explain something technical to someone who isn't.

For customer success, it's a different set: specific techniques for defusing an angry account, how they build trust early, how fast they pick up unfamiliar software, and examples of changing a customer's mind.

Write each competency as something you could point at in a transcript. Instead of "leadership potential", write "describes specific examples of mentoring team members or influencing a technical decision". If you can't imagine highlighting a sentence that proves it, it isn't a criterion—it's a vibe.

Step 2: Choose your rating scale

Five points, with definitions:

  • 1—Far below requirements. No evidence, avoided the question, or answered something else.
  • 2—Below requirements. Minimal understanding, vague, no concrete examples.
  • 3—Meets requirements. At least one relevant, complete example.
  • 4—Exceeds requirements. Multiple detailed examples or sophisticated reasoning.
  • 5—Far exceeds requirements. Exceptional mastery, non-obvious approaches.

Don't go past five points. Research on rating granularity is consistent that people can't reliably distinguish more than about five levels, so a 10-point scale buys you noise dressed as precision.

Step 3: Set up recording and transcription

Quality transcription starts with quality audio, and this is where most implementations lose more accuracy than any model choice could recover.

Platform setup. Zoom can record each participant to a separate audio file—turn that on, because separate channels beat any diarization model. Microsoft Teams and Google Meet give you a single mixed recording, which is fine but means you're relying on speaker separation.

Audio hygiene. External microphones over laptop mics, quiet rooms, consistent levels, wired connections where possible. Test before every interview, not after.

Consent. Tell candidates at scheduling that the interview will be recorded, get verbal consent on the call ("This interview will be recorded for evaluation purposes. Do you consent?"), and follow local law—several US states require all-party consent. Store recordings under access control and delete them on a defined schedule.

The reliability argument here isn't abstract. Metaview builds recruiting AI on AssemblyAI, and their CTO describes the difference this way:

"Since moving to AssemblyAI, we've seen a meaningful improvement in the confidence tail of our production transcripts....What stands out is not just the model quality, but the way [they] let us bring real meeting context into transcription, from calendar titles to organizations, domains, and participant names, so recruiting conversations come through with the nuance our customers depend on."

— Shahriar Tajbakhsh, Co-founder and CTO, Metaview

The confidence tail is the right thing to care about. Average accuracy is a comfortable number; what breaks a scorecard is the handful of low-confidence tokens that land on a technical term, a company name, or a negation.

Step 4: Configure AssemblyAI transcription in Python

Install:

pip install "assemblyai>=1.0.0" python-dotenv

Everything below runs against the async Speech-to-Text API. All code here targets Python SDK 1.0.0 (released August 14, 2026), which unified the async, streaming, and sync clients behind one shape. Store your key in a .env file, never in source.

Python:

import os

from dotenv import load_dotenv
from assemblyai import TranscriptStatus
from assemblyai.prerecorded.v2 import Transcriber, TranscriptionConfig

load_dotenv()

config = TranscriptionConfig(
    # Pin the model. Async takes speech_models — plural, an array.
    speech_models=["universal-3-5-pro"],

    speaker_labels=True,
    speakers_expected=2,          # interviewer + candidate

    # Prime the model with the vocabulary this role will actually use.
    keyterms_prompt=[
        "Kubernetes", "idempotency", "OpenTelemetry", "Sarah Okafor",
    ],

    punctuate=True,
    format_text=True,
)

transcriber = Transcriber(api_key=os.environ["ASSEMBLYAI_API_KEY"])
transcript = transcriber.transcribe(
    "interviews/okafor_backend_r2.m4a", config=config
)

if transcript.status == TranscriptStatus.error:
    raise RuntimeError(f"Transcription failed: {transcript.error}")

for utt in transcript.utterances:
    ts = f"{int(utt.start // 60000):02d}:{int(utt.start % 60000 // 1000):02d}"
    print(f"[{ts}] {utt.speaker}: {utt.text}")

Four things in that config are worth understanding rather than copying.

speech_models is plural and pinned. If an older tutorial—including an earlier version of this one—shows the deprecated singular speech_model parameter paired with the legacy best identifier, that line carries two deprecations at once. The singular parameter is deprecated for async, and the legacy identifier no longer maps to anything current. Requests using either shape re-route on September 2, 2026, which means your transcripts, your diarization output, and your per-hour cost all move on a date you didn't choose. Pin it.

speakers_expected matters more here than in almost any other use case, because every score in your scorecard is attributed to a speaker. Worth knowing: this parameter and the min/max bounds were being ignored on one diarization route until recently. That's been fixed, and advanced_speaker_segmentation now applies by default when you set it. If you built against this tutorial before, re-run it—the output may genuinely differ.

Keyterms prime the model on your vocabulary. Universal-3.5 Pro takes up to 1,000 terms. Load the role's technical stack and the interviewer's and candidate's names. This is the same mechanism as contextual prompting, and it's what keeps "idempotency" from becoming "a potency".

File formats aren't your problem. MP3, MP4, M4A, WAV—upload it as-is. No conversion step.

Gate low-confidence utterances before they become evidence

A misquoted or mis-attributed line isn't a cosmetic error. It's a wrong hiring decision with a paper trail that looks correct.

Python:

LOW_CONFIDENCE = 0.75

flagged = [u for u in transcript.utterances if u.confidence < LOW_CONFIDENCE]

if flagged:
    print(f"{len(flagged)} utterances need human review before scoring:")
    for u in flagged:
        print(f"  {u.speaker} ({u.confidence:.2f}): {u.text[:80]}...")

Every utterance carries a confidence score, and so does every word inside it. Route anything below your threshold to a human before it becomes evidence. That's a three-line safeguard on the highest-stakes output your pipeline produces.

Be clear about what it does and doesn't tell you. confidence is the model's certainty about the words, not a separate score for the speaker attribution—there is no per-utterance speaker-confidence field. In practice the two correlate: the utterances the model was least sure of are usually the crosstalk, the interruptions, and the moments two people talk at once, which are exactly where a turn gets assigned to the wrong person. Reviewing the low-confidence tail catches both failure modes at once, and it's a short list.

On the underlying quality: Universal-3.5 Pro produces the transcript and the speaker changes jointly rather than in two passes, and it's optimized for cpWER rather than DER—averaging 30.17 across meetings, telephony, far-field, and conversational audio, against Deepgram Nova-3 English at 37.92, ElevenLabs Scribe v2 at 35.26, and Gladia at 36.87. For raw word accuracy, Universal-3 Pro records a 5.6% mean English word error rate (median 4.9%) across our benchmark suite. Treat that as a floor rather than a promise: benchmark audio is clean, and an interview over a laptop mic in an untreated room will score worse. If you want the fuller picture, we've written up how accurate speech-to-text actually is in 2026.

Step 5: Extract evidence and calculate scores

Now the interesting part. For each competency, find the candidate's turns that contain evidence, keep the quote and the timestamp, and score on the quantity and depth of what you found.

Python:

from dataclasses import dataclass

COMPETENCIES = {
    "system_design": ["trade-off", "scale", "bottleneck", "latency", "partition"],
    "debugging":     ["root cause", "reproduce", "instrument", "hypothesis"],
    "collaboration": ["code review", "pair", "disagreed", "mentored"],
}

CANDIDATE = "B"   # confirm against your diarized output first

@dataclass
class Finding:
    competency: str
    quote: str
    timestamp: str
    confidence: float = 1.0

def timestamp(ms: int) -> str:
    return f"{int(ms // 60000):02d}:{int(ms % 60000 // 1000):02d}"

def gather_evidence(transcript, competencies, speaker):
    found = {name: [] for name in competencies}
    for utt in transcript.utterances:
        if utt.speaker != speaker:
            continue
        lowered = utt.text.lower()
        for name, cues in competencies.items():
            if any(cue in lowered for cue in cues):
                found[name].append(
                    Finding(name, utt.text, timestamp(utt.start), utt.confidence)
                )
    return found

def score(findings):
    # Depth proxy: distinct evidence moments, capped at the scale.
    return min(5, 1 + len(findings))

evidence = gather_evidence(transcript, COMPETENCIES, CANDIDATE)
scorecard = {name: (score(f), f) for name, f in evidence.items()}

Keyword matching is a starting point, not the destination. It's deliberately simple so you can see the shape, and it will produce false positives on any candidate who says "trade-off" while explaining why they avoid trade-offs. Use it to draft, then let a model do the judging.

Example scorecard output

Competency Score Evidence quote Timestamp
System design 4 "We partitioned by tenant rather than by time, which cost us on range queries but meant a noisy tenant couldn't take down the shard." 14:22
Debugging 5 "I couldn't reproduce it locally, so I added tracing around the retry path and found we were double-counting on the second attempt." 27:08
Collaboration 2 "Code review is mostly a formality on our team." 38:51

That table is the deliverable. A hiring manager can read it in fifteen seconds, and every row is falsifiable—jump to 38:51 and hear it in context.

Let a model do the judging

Keyword matching finds candidates for evidence. A language model decides whether the evidence is any good. Route that through LLM Gateway, which puts OpenAI, Anthropic, and Google models behind the same key you're already using:

Python:

import json, requests

candidate_turns = "\n".join(
    f"[{timestamp(u.start)}] {u.text}"
    for u in transcript.utterances if u.speaker == CANDIDATE
)

resp = requests.post(
    "https://llm-gateway.assemblyai.com/v1/chat/completions",
    headers={"authorization": os.environ["ASSEMBLYAI_API_KEY"]},   # no Bearer
    json={
        # Pin a current model id — the gateway has no "-latest" aliases,
        # and every model carries a published retirement date.
        "model": "claude-sonnet-5",   # or claude-sonnet-4-6, gemini-3.5-flash
        "max_tokens": 1000,
        "messages": [{
            "role": "user",
            "content": (
                "Score this candidate 1-5 on system design, debugging, and "
                "collaboration. For each, return the single strongest verbatim "
                "quote and its timestamp. If there is no evidence, score 1 and "
                "say so. Return JSON only.\n\n" + candidate_turns
            ),
        }],
    },
    timeout=90,
)
print(json.dumps(resp.json()["choices"][0]["message"]["content"], indent=2))

Two rules for this step. Constrain the model to quoting rather than paraphrasing—a summarized "showed strong debugging instincts" is exactly the unauditable judgement you built this system to avoid. And require it to say when there's no evidence, because a model asked to score will always produce a number.

Validate Transcription Quality On Your Audio

Upload an interview clip and check accuracy, speaker labels, and timestamps in your browser. No code required before you wire up the API.

Try playground

Handling candidate data

Interview recordings are personal data about people who don't work for you and may never will. Treat them accordingly.

Use PII redaction to strip names, addresses, and identifiers from stored transcripts. You can redact in text and in audio, which means you can retain a recording for QA without retaining the candidate's phone number. Set a retention window tied to the hiring decision, not to "forever", and restrict transcript access to the panel.

If your data has to stay in the EU, api.eu.assemblyai.com is the same price as the US endpoint.

Three mistakes that break scoring systems

One-size-fits-all criteria

The same competency list for every role guarantees the list is generic, and generic criteria produce scores that cluster in the middle. Analyze your actual top performers per role. If your best engineers are the ones who give useful code review, score "provides constructive technical feedback", not "teamwork".

Skipping calibration

Transcripts don't fix disagreement, they just make it visible. Run a monthly session where every interviewer scores the same sample transcript independently, then compare. When two people differ by more than a point, the interesting question is never who was right—it's which unstated assumption they're each carrying.

Neglecting audio quality

Poor audio ruins everything downstream. A transcript with 20% errors can invert the meaning of a technical explanation, and the score you produce from it will look every bit as confident as a correct one. Test the recording setup before each interview, require external mics, use wired connections, and be willing to reschedule when the audio is bad. One garbled explanation can move a candidate from "exceeds" to "below".

Measuring whether the system works

Four metrics, tracked before and after you roll this out:

  • Time-to-hire. Days from posting to accepted offer. Clear evidence should shorten debates.
  • Quality of hire. Performance ratings at six months, correlated against interview scores. If there's no correlation, your criteria are measuring the wrong thing.
  • Inter-rater reliability. Agreement between evaluators on the same interview. Cohen's Kappa above 0.7 is generally considered good agreement.
  • Score distribution. If everyone lands on 3, your criteria are too generic or your anchors are too soft.

Python:

from sklearn.metrics import cohen_kappa_score

rater_a = [4, 3, 5, 2, 4, 3]
rater_b = [4, 4, 5, 2, 3, 3]

kappa = cohen_kappa_score(rater_a, rater_b, weights="linear")
print(f"Weighted kappa: {kappa:.2f}")   # > 0.70 = good agreement

Technical hiring platforms including HackerRank build on AssemblyAI. If you're evaluating whether to build this or buy it, we've mapped the wider category in our piece on hiring intelligence platforms, and the same transcript infrastructure underpins conversation intelligence solutions elsewhere in the business.

Interview scoring sheet vs AI scoring system

Most teams searching for this start with a template—a PDF or a spreadsheet with competencies down the side and a 1–5 scale across the top. That's a scoring sheet, and it's genuinely useful: it's the artifact that forces you to define criteria before the interview instead of rationalizing after it.

An AI scoring system doesn't replace the sheet. It fills it in with evidence. Keep the rubric exactly as it is, and change only where the score comes from—memory becomes transcript, and "strong answer on system design" becomes a quote at 14:22.

Which means the honest sequencing advice is: build the sheet first. If your competencies aren't crisp, automating the scoring just produces confident nonsense faster.

What about interviewing with an agent?

There's a version of this where the interview itself is conducted by software—structured screening rounds, consistent question ordering, no scheduling. AssemblyAI's Voice Agent API supports that shape: one WebSocket for speech-to-text, reasoning, and speech output at a flat $4.50 per hour, with server-side tool calling and a session history endpoint that returns the recording and transcript when the session ends.

That last part is what makes it relevant here. An agent-run screen drops into exactly the same scoring pipeline you just built—same transcript, same evidence extraction, same scorecard. The interview method changes; the evaluation layer doesn't.

Whether you should is a separate conversation, and a candidate-experience one rather than a technical one.

What this actually changes

The obvious pitch for transcript-based scoring is consistency, and consistency is real. But it's not the biggest thing.

The biggest thing is that scoring stops being a claim and starts being a citation. When a rating carries a quote and a timestamp, disagreement between interviewers becomes productive instead of political—you're no longer arguing about whose memory is better, you're both looking at 38:51. That's a different conversation, and it's the one where hiring bars actually get calibrated.

A 45-minute interview costs about $0.17 all-in to transcribe with speaker labels: Universal-3.5 Pro at $0.21 per hour plus standard diarization at $0.02 per hour, billed per second with no minimums. LLM Gateway tokens for the scoring pass are separate. Against the interviewer time a single bad hire consumes, that's a rounding error—which is the strange part. The cheapest component in your hiring process is the one that makes the whole thing defensible.

Start Scoring From Evidence, Not Memory

Create an account to access the Speech-to-Text API with speaker diarization and word-level timestamps, then plug it into the scoring workflow above.

Sign up free

Frequently asked questions

How do interviewers score you?

Most structured processes rate each candidate against three to seven predefined competencies on a fixed scale, usually 1-5, with a written behavioral anchor describing what each score looks like. Interviewers score during or immediately after the conversation and record supporting notes. The weakness is recall: notes are written from memory, so the evidence behind a score is a paraphrase rather than what the candidate actually said.

What are the 5 C's of interviewing?

The 5 C's are commonly given as competence, character, communication, culture fit, and career direction—a memory aid for the dimensions a structured interview should cover. They are a starting point, not a rubric: each one needs a role-specific behavioral definition and a scale before it can be scored consistently across a panel.

How accurate does speech-to-text need to be for reliable interview scoring?

Accurate enough that the evidence quotes attached to each score are verbatim, because a scorecard's whole value is that a disputed rating can be re-read. Universal-3 Pro records a 5.6% mean English word error rate across our benchmark suite, but interview audio recorded on a laptop mic in an untreated room will score worse—microphone quality moves accuracy more than model choice does at this level.

Can I use this with video interview platforms like Zoom and Teams?

Yes. Both Zoom and Microsoft Teams can save cloud recordings as audio or video files, and async transcription accepts either by URL or upload. Point the pipeline at the recording location, transcribe with speaker labels so interviewer and candidate turns separate cleanly, and run the scoring pass over the candidate's turns only.

How much does it cost to transcribe and score an interview?

A 45-minute interview costs about $0.17 all-in to transcribe with speaker labels—Universal-3.5 Pro at $0.21/hr plus standard diarization at $0.02/hr—billed per second with no minimums. The LLM scoring pass through LLM Gateway is billed separately on tokens. For most hiring teams the transcription cost per hire is a rounding error against the interviewer time it recovers.

What should I do if a candidate refuses to consent to interview recording?

Run the interview unrecorded and score it manually against the same rubric—consent to recording must never be a condition of consideration, and in two-party-consent jurisdictions recording without it is unlawful. Design the process so the rubric, not the recording, is the constant: the transcript improves the evidence behind a score, it does not define the criteria.

‍

Title goes here

Lorem ipsum dolor sit amet, consectetur adipiscing elit, sed do eiusmod tempor incididunt ut labore et dolore magna aliqua. Ut enim ad minim veniam, quis nostrud exercitation ullamco laboris nisi ut aliquip ex ea commodo consequat. Duis aute irure dolor in reprehenderit in voluptate velit esse cillum dolore eu fugiat nulla pariatur.

Button Text
HR