Speech-to-text for HR and recruiting: Interview transcription, screening and scoring
Interview transcription software for HR records, screens, and scores candidates with speaker labels, searchable transcripts, and compliance support.



Here's a failure mode every recruiting team has hit. A hiring manager searches the transcript archive for a candidate, finds nothing, and concludes the interview was never recorded. It was. The model wrote her name three different ways across forty minutes, and none of them matched the ATS.
Interview transcription isn't general transcription with a recruiting label on it. The words carrying the most weight—candidate names, companies, universities, the specific technology someone claims to have shipped—are exactly the words general speech models get wrong. And in hiring, a wrong transcript isn't a typo. It's an evaluation record you may have to defend.
What is interview transcription software?
Interview transcription software converts recorded or live interviews into searchable, attributed text—speaker labels, timestamps, and structured output for scoring or ATS ingestion. In practice it's a speech-to-text API plus whatever your team builds around it, which is why most serious recruiting products build rather than buy.
Recruiting needs four things from it: multi-speaker attribution on panels, name accuracy good enough to search on, redaction before anything hits storage, and structured output that maps to a scorecard. Notetakers hand you a wall of text. A recruiting product needs the fields.
Why interview accuracy comes down to names
Universal-3.5 Pro averages 4.35% normalized word error rate across our async benchmark datasets. That's the headline number, and it's the least interesting one for hiring.
What matters in recruiting is the name entity error rate—how often the model mangles a proper noun. On AssemblyAI's English voice-agent benchmark of live, conversational audio, the gap between providers there is wide:
| Entity type | Universal-3.6 Pro Realtime | ElevenLabs Scribe v2 | Deepgram Flux EN |
|---|---|---|---|
| Names | 10.9% | 14.8% | 29.0% |
| Overall entity error rate | 14.4% | 18.5% | 30.1% |
| Word error rate | 5.19% | 7.78% | 13.50% |
Look at the spread between rows. Word error rate varies by several points across providers. Name error rate varies by nearly twenty, and even the closest competitor misses names about a third more often. Evaluate an interview stack on WER alone and you're grading it on the dimension where vendors look most alike, ignoring the one where they don't.
How do you get a model to know a candidate's name in advance?
You tell it. Contextual prompting primes the model before transcription with what's already sitting in the calendar invite and the ATS record—the candidate's name, the company, the role title, the panel, the domain names in the meeting link.
This is the single biggest win most recruiting products can pick up, because the information is already in your system. You're not asking the model to guess an unusual surname from audio. You're handing it a shortlist.
Metaview builds AI notes for recruiters—interview transcription is the product, not a feature of it:
"Since moving to AssemblyAI, we've seen a meaningful improvement in the confidence tail of our production transcripts....What stands out is not just the model quality, but the way [they] let us bring real meeting context into transcription, from calendar titles to organizations, domains, and participant names, so recruiting conversations come through with the nuance our customers depend on." — Shahriar Tajbakhsh, Co-founder and CTO, Metaview
The "confidence tail" phrasing is worth sitting with. Low-confidence tokens are the ones that force a human to open the audio and check, and that tail is what thins out on Universal-3.5 Pro for async. Fewer trips back to the recording changes the economics of a review workflow more than a fractional WER improvement ever does.
HackerRank runs technical assessment and coding interviews—where the vocabulary problem is framework names and algorithm terminology rather than surnames, but the fix is the same.
Drop in a real interview recording, add the candidate and panel names as context, and compare the two transcripts side by side.
Can speaker diarization handle a panel interview?
Yes, and panels are the hard case. Four people, cross-talk, a candidate answering before the question finishes, one-word interjections a clustering pass will happily assign to whoever was loudest.
Universal-3.5 Pro produces the transcript and the speaker changes jointly, in one model, rather than transcribing first and running a separate speaker-clustering step over the result. That's the architectural reason it holds up on short turns, rapid back-and-forth, and overlapped speech—the exact conditions of a panel interview.
It's optimized for cpWER rather than DER, and that distinction matters here. DER asks whether the speaker boundaries landed in the right place; cpWER asks whether the right words ended up attributed to the right person—which is what a hiring manager reading a transcript needs.
| Model | cpWER (lower is better) |
|---|---|
| Universal-3.5 Pro | 30.17 |
| ElevenLabs Scribe v2 | 35.26 |
| Gladia | 36.87 |
| Deepgram Nova-3 English | 37.92 |
On live interviews, streaming speaker diarization labels speakers as the conversation happens, then re-clusters and sends one correction within about half a second of the stream closing. Your live UI stays responsive; the stored transcript is the corrected one. For the mechanics, see how speaker diarization works.
Voice screening agents are the fastest-growing use case in recruiting
Interview agents and lead qualification are the number one ranked use case for the Voice Agent API—not a side application, the largest single category of what people build on it.
The reason is structural. First-round screening is high-volume, scripted, and almost entirely about collecting a fixed set of facts: work authorization, notice period, compensation range, location, a few role-specific qualifiers. It's the part of recruiting that scales worst with headcount and where a five-day delay costs you candidates.
What makes it buildable now rather than a six-month integration project:
- One WebSocket, not three vendors. Speech-to-text, LLM, and text-to-speech behind a single connection. One bill, one set of logs, one place to look when a call goes wrong.
- Flat $4.50/hr, all in. STT plus LLM plus TTS. Price a screening program off one number instead of modeling three usage curves. Full breakdown on the pricing page.
- ~1–1.3 seconds end-to-end. Fast enough that candidates don't talk over the agent, which is where screening calls usually fall apart.
- No SDK required. Standard JSON over a WebSocket. If your stack can hold a socket open, it can run a screening agent.
- Tool-parameter guardrails. Screening agents write into systems—booking a next round, updating a candidate record. The API constrains tool arguments to what is actually inferable from the candidate's answers and from prior tool results, rather than letting the model fill a field with a plausible-looking value.
That last point decides whether a screening agent is a demo or a production system. An agent that occasionally misreports a candidate's availability isn't a scheduling assistant, it's a source of recruiter cleanup work. If you're scoping a build, the build guide covers the connection shape, the tool-calling shape, and the single-socket architecture.
STT, LLM, and TTS at a flat $4.50/hr with no SDK to install. Start free and make your first call today.
How do you keep candidate PII out of the transcript?
Redact it before it lands. Interview audio is dense with exactly the data an HR team least wants sitting in a general-purpose datastore: full names, home addresses, phone numbers, dates of birth, sometimes salary history in jurisdictions where asking is restricted.
PII redaction through the Speech Understanding API runs across both the transcript text and the audio itself, so the stored recording has identifiers bleeped too. You can keep an evaluation record without keeping the identifiers that make it a liability. Guardrails handles the other direction—content moderation and safety on what goes in and comes out—which stops being optional the moment an agent is talking to candidates unsupervised.
How do interview transcripts get into your ATS?
Two mechanisms, and together they cover most of what a Greenhouse, Lever, Workday, or BambooHR workflow needs.
Webhook callbacks handle the timing problem. Submit the interview audio, get an ID back immediately, register a callback URL. When transcription finishes we POST to your endpoint, and your service writes the result to the candidate record. No polling loop, no worker on a timer, no interviews stuck in a queue because a cron job died at 2am.
The Speech Understanding request object handles the shape problem. Rather than dumping text on the ATS and asking a downstream job to parse it, request the structured pieces in the same call: entity detection across 50+ types, sentiment, topics, key phrases, and redaction. What comes back maps onto scorecard fields. Request and response shapes are in the API reference.
This beats a pre-built connector for one reason: every company's competency framework is different, and a fixed integration gives you a transcript field and nothing else. Structured output plus a webhook gives you the fields you score on.
Which model should you use for interviews?
For recorded interviews, Universal-3.5 Pro at $0.21/hr—the default async model for all accounts since August 2026, covering 18 languages with native code-switching and no separate language pass. For live interviews and screening agents, Universal-3.6 Pro Realtime at $0.45/hr base, covering 32 languages, billed on session duration rather than audio duration.
One correction worth stating plainly, because the wrong version circulates: the Universal Pro line is 18 languages for Universal-3.5 Pro and 32 for Universal-3.6 Pro Realtime. The 99+ figure belongs to Universal-2, which stays available at $0.15/hr and is the documented fallback for coverage beyond the Pro line.
Final words
Most teams run a bake-off on word error rate, pick the winner, then spend the next quarter fielding complaints about names.
Run it differently. Take twenty real interviews, pull out every proper noun—candidate names, companies, schools, technologies—and score only on those. Then re-run the same audio with the candidate and panel names supplied as context, and score again. The delta between those two runs is worth more than the WER column, and it's the one thing a generic benchmark will never tell you about your own pipeline.
The four-person panel with the cross-talk and the hard-to-spell surname. That's the one that tells you whether a model is worth building on.
Frequently asked questions
What is the most reliable speech-to-text API for interview transcription?
For recorded interviews, Universal-3.5 Pro averages 4.35% normalized word error rate. For live interviews, Universal-3.6 Pro Realtime has the lowest name entity error rate of the models we benchmark—10.9% against 14.8% for ElevenLabs Scribe v2. Name accuracy is the metric that decides whether an interview archive is searchable, so it's a better selection criterion for hiring than word error rate alone.
Is it possible to transcribe multi-speaker interviews automatically?
Yes. Universal-3.5 Pro produces the transcript and the speaker changes together in one model rather than running a separate clustering pass, which is why it holds up on the short turns and overlapped speech typical of panel interviews. It's optimized for cpWER, scoring 30.17 against 37.92 for Deepgram Nova-3 English. On live interviews, speakers are labeled in real time and then corrected once within about half a second of the stream ending.
How do I use APIs to highlight named entities in interview transcripts?
Entity detection is part of the Speech Understanding request object and covers 50+ entity types, so you can pull out names, organizations, locations, and dates in the same call that produces the transcript. Combined with a webhook callback, that gives you structured fields ready to write into a candidate record rather than raw text a downstream job has to parse. The same request also handles PII redaction across both text and audio.
Which speech-to-text provider do production voice agents rely on?
Interview agents and lead qualification are the top-ranked use case on our Voice Agent API, which runs on Universal-3.6 Pro Realtime as its speech foundation. On AssemblyAI's English voice-agent benchmark it records a 5.19% word error rate and a 14.4% entity error rate, both the best in that comparison set. Tool-parameter guardrails constrain what an agent can pass to your systems, limiting arguments to what is inferable from the candidate's answers and from prior tool results.
What is AssemblyAI's Voice Agent API and how does it work?
It's a single WebSocket that handles speech-to-text, the LLM, and text-to-speech behind one connection, at a flat $4.50/hr covering all three. End-to-end latency runs around 1 to 1.3 seconds, it speaks standard JSON with no SDK required, and it ships drop-in plugins for LiveKit and Pipecat. For recruiting, that means a screening agent is a socket connection and a tool schema rather than a three-vendor integration.
How does interview transcription handle candidate PII and data retention?
PII redaction runs across both the transcript text and the audio file, so stored recordings have identifiers removed rather than just the text copy. That lets you retain an evaluation record for the period your policy requires without retaining the personal data that makes it a liability. Guardrails adds content moderation and safety on top, which matters most for screening agents talking to candidates without a human on the line.
Lorem ipsum dolor sit amet, consectetur adipiscing elit, sed do eiusmod tempor incididunt ut labore et dolore magna aliqua. Ut enim ad minim veniam, quis nostrud exercitation ullamco laboris nisi ut aliquip ex ea commodo consequat. Duis aute irure dolor in reprehenderit in voluptate velit esse cillum dolore eu fugiat nulla pariatur.



