How to evaluate a speech-to-text API for an education platform (2026)
How to evaluate a speech-to-text API for an education platform: student-speech accuracy, archive economics, FERPA, accessibility, and a buyer scorecard.



There's a generic version of this evaluation, and we've written it: how to choose the best speech-to-text API for your product covers accuracy, latency, languages, pricing, scale, compliance, developer experience, and support. If you're starting cold, start there.
This post is the education-specific layer on top, because that generic checklist quietly assumes a few things that aren't true for education platforms. It assumes your audio is roughly uniform. It assumes your volume is roughly steady. It assumes your compliance question is SOC 2 and maybe HIPAA. And it assumes the person speaking is the person your accuracy matters for.
None of those hold when you're transcribing a lecture hall, a tutoring session, or an eight-year-old reading aloud.
Here's what actually decides this evaluation.
1. Your accuracy problem is students, not instructors
This is the biggest blind spot, and almost every vendor evaluation gets it backwards.
Instructor audio is the easy case. A professor lecturing is doing the thing speech models are best at: a single speaker, speaking deliberately, in complete sentences, usually with a decent microphone, on a topic with consistent vocabulary. Every vendor will look fine on a lecture recording. If you evaluate on lecture audio alone, you will learn nothing that distinguishes them.
Student audio is a different problem entirely, and it's where your product lives or dies:
- Accented and non-native speech. International students, ESL learners, and regional accents that are underrepresented in training data. If your platform serves a global student body, this is your dominant accuracy risk, and it's the one least visible in vendor benchmarks.
- Children's speech. Genuinely hard, and categorically different from adult speech — higher pitch, different phonetics, incomplete articulation, unpredictable pacing. Models that are excellent on adult audio can be unusable here. If you're building for K-12 or early literacy, this is the entire evaluation.
- Classroom acoustics. A student asking a question from the back of a lecture hall, into no microphone, over HVAC noise. Far-field, low SNR, and frequently overlapping with other speech.
- Code-switching. Multilingual students switching mid-sentence, which a model committed to one language handles by producing nothing or producing nonsense.
What to do: build your evaluation set from your own student audio, weighted toward your hardest cohorts, not from clean lecture recordings. Include at least one file per accent group you serve, at least one far-field question from a real room, and if you serve young learners, make children's speech the majority of the set. Then read how to evaluate speech recognition models for the methodology and why your WER benchmark might be lying to you before you trust the resulting numbers.
This is not a hypothetical differentiator. An AI reading tutor for children evaluated the market over several months and concluded that only one streaming model was accurate enough for children's speech — the hardest audio in the category — and that conclusion, not price, decided the deal. If your hardest cohort is the one your product is for, it should be the one your evaluation is built around.
Upload accented, far-field, or young-learner recordings and compare transcripts directly. Clean lecture audio won't tell you what you need to know.
2. The economics of long-form audio, at catalog scale
Education audio is long. A lecture is 50 to 90 minutes. A course is 30 lectures. A catalog is thousands of courses. Per-minute pricing that feels trivial on a 4-minute support call compounds very differently here.
Model the two workloads separately, because they behave nothing alike:
| Workload | Shape | What it breaks |
|---|---|---|
| Archive backfill | One-time, enormous, latency-insensitive | Budgets, and throughput ceilings |
| Ongoing capture | Recurring, predictable, extremely seasonal | Capacity planning, and any commit-based contract |
| Live captioning | Concurrent, latency-critical, bursty at class times | New-session rate limits, at the start of every class hour |
Three things to press vendors on:
Seasonality. Academic volume is not smooth. It collapses over summer and spikes hard at the start of term — the back-to-school ramp is real and it is steep. Any pricing model built on annual commits will either overcharge you for eight months or throttle you in September. Pay-as-you-go with no minimums matches the shape of the business; AssemblyAI has no contracts or minimums and bills pre-recorded audio to the exact second, which means the summer trough costs you nothing.
Throughput during backfill. If you're processing a ten-thousand-hour archive, the question isn't price per hour — it's how fast you can push it through, and what happens when you push harder than the vendor expects.
Get the specific mechanism rather than a reassurance. For AssemblyAI pre-recorded transcription, it's this: your account has a limit on how many transcription jobs run in parallel — 5 on free accounts, 200+ on paid — and anything you submit beyond that is queued FIFO and processed automatically as earlier jobs finish, not rejected. So over-submitting costs you latency, not lost work, and a backfill is a throughput-planning exercise rather than an error-handling one. Higher limits are available on request at no additional cost, which is the thing to arrange before you start a ten-thousand-hour job rather than during it.
Two related ceilings worth knowing: there's a separate HTTP limit of 20,000 requests per five minutes across all endpoints, which you hit by polling too eagerly rather than by transcribing too much — use webhooks, or widen and jitter your polling. And if your account balance goes negative, the parallel limit drops to 1, which is a memorable way to discover a billing problem mid-backfill.
The accuracy-to-cost tradeoff is really an accuracy-to-labor tradeoff. This is the part finance teams miss. If your captions need human review before publication — and for accessibility they generally do — then a few points of accuracy translate directly into review hours. As the CTO of one video platform put it:
"If you have an hour of content, the difference between 99% accuracy and 97% accuracy, it's a lot of time for that person to review. So you could cut down their workflow from taking half an hour, taking 20 minutes, taking 15 minutes — it's huge, right?"
— Joshua Grossberg, CTO, Kapwing
Run that math on your own catalog before you optimize for the cheapest per-hour rate. At scale, review labor usually dominates transcription spend, and the cheaper model is frequently the more expensive choice.
For reference, AssemblyAI's pre-recorded pricing is $0.21/hr for Universal-3.5 Pro and $0.15/hr for Universal-2 where you need coverage across 99 languages. Streaming rates, including Universal-3.6 Pro, are on the pricing page.
3. Student data is its own compliance question
Education compliance is not healthcare compliance with different words. It's a distinct set of obligations, and vendors who lead with SOC 2 and HIPAA often can't answer the questions that actually block an edtech deal.
Student education records. In the US, FERPA governs them, and the obligations flow to you as the platform — which means they flow to your vendor through your contract with them. The practical questions are: is audio retained, for how long, is it used for model training, and can you turn that off. Get the answers in writing before you build.
Age restrictions, including the indirect kind. This one catches teams off guard. It's not just whether your users are old enough to accept terms of service — it's whether the content of the audio concerns someone who isn't. A K-12 product where teachers, not students, do the recording still produces recordings that contain information about minors. If your product touches K-12 at all, raise this with your vendor and your own legal team early, because it can require contract terms that don't exist in a standard signup flow.
Institutional data agreements. Universities increasingly require a Data Security Agreement or DPA before they'll let a department keep using a service, and these arrive without warning from procurement offices that weren't part of your sales conversation. Ask vendors whether they offer a standard DPA and whether they'll provide terms in an editable format. A vendor who can't will cost you months.
Data residency. European institutions will ask where audio is processed. AssemblyAI offers EU data residency through api.eu.assemblyai.com for pre-recorded audio and streaming.eu.assemblyai.com for streaming.
4. Accessibility is a procurement gate, not a feature
For higher education and public-sector buyers, captions are a legal obligation under ADA and Section 508, and the bar is set by what's defensible, not by what's impressive in a demo.
Which means two things for your evaluation:
There is an accuracy threshold, and you should find out who set yours. Institutional captioning policies commonly cite a 99% accuracy bar, but that figure comes from institutional policy and procurement practice rather than from the statutes themselves, and different institutions apply it differently. Ask the specific institution you're selling into what standard they enforce and how they measure it, rather than assuming a number. Then answer the empirical question — whether a given model clears that bar on your audio — with your own files, because that's the part your general counsel will ask about.
Caption mechanics matter more than they look. Getting WebVTT or SRT out of an API is easy. Getting it right at scale is where the real problems live — cue-level timing that stays aligned, speaker attribution in the caption track, and, if you translate, per-cue translation that preserves the original timing. That last one is a genuine trap: workflows that translate at the utterance or speaker-turn level rather than cue by cue produce badly misaligned subtitles in every target language, and teams usually discover it after they've built the pipeline. If you're shipping captions in multiple languages, test the timing alignment on a long file before you commit to an architecture.
Ask about a VPAT. Higher-ed and public-sector procurement routinely requires an Accessibility Conformance Report. Ask every vendor on your list whether they have one. Their answer, and how quickly they give it, tells you something about how many education deals they've actually closed.
Universal-3.5 Pro at $0.21/hr, pay-as-you-go with no minimums — built for backfilling archives and handling the September spike.
5. Build, buy, or self-host
Institutions with GPU budgets and research staff will reasonably ask why they shouldn't self-host an open model. It's a fair question and the answer isn't automatic.
Self-hosting makes sense when you have genuine data residency constraints that no vendor satisfies, you have staff who will own the pipeline for years rather than for a semester, and your accuracy requirements are met by an available open model on your actual audio.
It usually doesn't when your workload is seasonal — you're buying GPU capacity for a September peak and idling it through July — or when your hardest audio is the audio open models handle worst, which for education means children's speech, heavy accents, and far-field classroom questions. The gap between a self-hosted open model and a current commercial model is widest exactly where education needs it to be narrowest.
The honest middle path: many platforms use a commercial API for the audio that's hard and hosted infrastructure for the audio that's easy. If you're going to model this, model it against your real cohort mix rather than an average.
The scorecard
Take this into your evaluation:
| Criterion | The question to ask | Disqualifying answer |
|---|---|---|
| Student-speech accuracy | WER on our accented, young-learner, and far-field audio | Only benchmarks on clean read speech |
| Multilingual + code-switching | Handled natively, or as a separate detection pass? | Requires committing to one language per file |
| Diarization | Does it survive short student questions over instructor speech? | Post-hoc clustering that smears short turns |
| Long-file handling | 90-minute files without chunking or quality drift? | Duration caps that force you to split lectures |
| Backfill throughput | What's the parallel-job limit, what happens past it, and can it be raised? | Over-limit submissions are rejected rather than queued, or the limit is fixed |
| Live-captioning capacity | What's the new-session rate limit when 40 classes start at once? | No answer, or a fixed ceiling you can't raise |
| Seasonality fit | What do we pay in July? | Annual commits or minimums |
| Caption output | WebVTT/SRT with speaker labels; per-cue translation timing? | Translation at utterance level — misaligns every language |
| Student data | Retention, training opt-out, DPA availability, age terms | No editable DPA; can't answer the K-12 question |
| Accessibility | VPAT/ACR available? Accuracy against the institution's bar on our audio? | No answer on either |
| Support | Who do we reach at 8am on the first day of term? | Ticket queue only |
The thing worth deciding first
Most education platform teams run this evaluation on lecture audio because lecture audio is what they have lying around. Every vendor scores well, the differences look like rounding errors, and the decision comes down to price.
Then the product ships and the complaints are all about the student who asked a question from the back of the room, the international student whose name is transcribed four different ways, and the second-grader the reading app can't understand.
Those were the cases that should have decided the evaluation. They're the reason the differences between vendors exist at all — on easy audio, every modern model is fine. Build your test set out of your hardest cohort and the decision usually makes itself.
Free API key, pay-as-you-go pricing with no commitments, and parallel-job limits that scale to any backfill on request. Start with your hardest files.
Frequently asked questions
What is the best speech-to-text API for an education platform?
There isn't a single answer, because the decision is determined by your hardest student cohort rather than by average benchmark accuracy — every modern model performs well on clean instructor audio. Build your evaluation set from accented, young-learner, and far-field classroom recordings, then weigh long-file handling, backfill throughput, seasonal pricing fit, caption output quality, and student-data terms. AssemblyAI's Universal-3.5 Pro runs at $0.21/hr pay-as-you-go with no minimums, with native code-switching across 18 languages.
How much does it cost to transcribe a university's lecture archive?
Model backfill and ongoing capture separately, since they have completely different shapes. At $0.21/hr for pre-recorded transcription, a 10,000-hour archive is roughly $2,100 in transcription cost — but the binding constraint is usually throughput rather than price, so ask vendors how many jobs run in parallel and what happens when you exceed that. With AssemblyAI the parallel limit is 200+ on paid accounts and over-limit submissions are queued FIFO rather than rejected, with higher limits available on request at no extra cost. Factor in human review labor too: at accessibility-grade accuracy requirements, review hours frequently cost more than transcription does, which is why a cheaper, less accurate model is often the more expensive choice.
What accuracy do captions need for ADA and Section 508 compliance?
Institutional captioning policies commonly cite a 99% accuracy bar, but that figure comes from institutional policy and procurement practice rather than from the statutes themselves, and different institutions enforce it differently — so ask the institution you're selling into what standard they apply and how they measure it. Whether a given model clears that bar on your specific audio is then an empirical question you answer with your own files, particularly for student speech and far-field classroom questions. Higher-ed and public-sector procurement also commonly requires a VPAT or Accessibility Conformance Report from vendors, so ask for one early.
Does FERPA apply to speech-to-text vendors?
FERPA obligations sit with the educational institution and the platform serving it, and they reach your transcription vendor through your contract rather than directly. The practical questions to settle in writing are whether audio is retained, for how long, whether it's used for model training, and whether you can disable that. Institutions increasingly also require a Data Security Agreement or DPA before approving continued use, so confirm your vendor offers one in an editable format.
Can speech-to-text APIs handle children's speech?
Some can, and the variance between models is much larger here than on adult speech. Children's audio differs categorically — higher pitch, different phonetics, incomplete articulation, unpredictable pacing — so a model that performs excellently on adult benchmarks can be unusable for early-literacy products. If you're building for K-12 or reading instruction, make children's speech the majority of your evaluation set rather than a spot check, because it will be the deciding factor.
Should we self-host an open speech model instead of using an API?
Self-hosting makes sense when you have data residency constraints no vendor satisfies and staff who will own the pipeline for years. It usually doesn't for education, for two reasons: academic volume is highly seasonal, so you're buying GPU capacity for a September peak and idling it through July, and the accuracy gap between open models and current commercial models is widest on exactly the audio education needs most — children's speech, heavy accents, and far-field classroom questions. Many platforms land on a hybrid, using a commercial API for the hard audio.
Lorem ipsum dolor sit amet, consectetur adipiscing elit, sed do eiusmod tempor incididunt ut labore et dolore magna aliqua. Ut enim ad minim veniam, quis nostrud exercitation ullamco laboris nisi ut aliquip ex ea commodo consequat. Duis aute irure dolor in reprehenderit in voluptate velit esse cillum dolore eu fugiat nulla pariatur.



