Speech-to-Text for EdTech: Lectures, Captions, Accessibility & Assessments
Lecture transcription helps students turn class recordings into searchable text for studying, captions, and accessibility. Learn recording tips and laws.



Most articles about lecture transcription treat it as a study hack. It isn't, or at least that's not where it starts.
For a deaf student, a transcript is the lecture. For a student with an auditory processing disorder, it's the difference between following an argument and reconstructing it afterwards from fragments. And for the institution running the lecture hall, captioning is an obligation, not a feature request.
So let's start there, and get to the study workflows after.
Captioning lectures is an accessibility obligation, not a nice-to-have
Public universities and institutions receiving federal funding in the US have legal duties around effective communication and accessible electronic content, under the Americans with Disabilities Act and Section 508 of the Rehabilitation Act. In practice that means recorded lecture content needs captions or transcripts, and live instruction needs a real-time path for students who can't rely on audio.
This is why lecture transcription tends to arrive as an institutional project rather than a student one. Someone in disability services or the online learning office is looking at a library of thousands of recorded hours and working out what it costs to make all of it accessible. That's a very different problem from one student wanting their notes searchable — and it's the problem speech-to-text infrastructure is actually good at.
What transcripts do for students with disabilities
Hearing impairment and auditory processing disorders
For deaf and hard-of-hearing students, a transcript delivers the full content of a lecture in a form they can read at their own pace, including the parts a lipreader or a partial hearing aid signal would miss — the professor turning to the whiteboard mid-sentence, the question shouted from the back row.
Auditory processing disorder is a different problem with a similar solution. The audio comes through fine; assembling it into meaning in real time is what fails. A transcript moves that work from the auditory channel to the visual one, where it isn't bottlenecked.
ADHD and executive function
Attention doesn't fail evenly across 50 minutes. It fails in patches, and the patches are unpredictable. A student with ADHD can be fully engaged for 40 minutes and lose the eight minutes where the professor explained the thing the exam is about.
A transcript makes those gaps recoverable without replaying the entire recording. Searching for a term and landing on the two paragraphs where it was defined removes both the time cost and the low-grade anxiety of not knowing what you missed. It also decouples note-taking from listening, which is worth a lot when splitting attention between the two is the specific thing that's hard.
How to request transcription as an accommodation
The process is broadly consistent across US institutions. Register with your disability services office — the name varies, "accessibility services" and "student accessibility" are common. Provide documentation from a qualified professional. Then ask specifically for lecture transcription or captioning, in writing, rather than a general accommodation.
Two things worth knowing. Accommodations can be revised: if what you're granted turns out to be insufficient, you can go back. And if the institution declines, students routinely arrange their own recording and transcription under the same accommodation that permits recording — which is where the rest of this post becomes relevant.
Upload a recording and get a searchable, speaker-labelled transcript back in a couple of minutes. Free to start, no credit card.
Live captioning in the lecture hall
Transcripts after the fact solve review. They don't solve being in the room. For that you need real-time transcription — audio streaming over a WebSocket, text appearing on a laptop or a projected caption feed as the lecturer speaks.
Universal-3.6 Pro Realtime runs at $0.45/hr, billed on how long the session stays open rather than how much audio you send. Its end-of-turn detection weighs whether what has been said so far reads as a completed thought rather than waiting out a fixed silence — which matters in a lecture, where a professor pausing to write on a board isn't finishing a sentence.
It also labels speakers live and then re-clusters and sends a single correction within about half a second of the stream ending, up to 10 speakers. In a seminar with open discussion, that's the difference between a caption feed and a usable one.
How accurate is lecture transcription?
Accurate enough to study from, and the failure modes are predictable. Universal-3.5 Pro averages 4.35% normalized word error rate across our evaluation datasets — the full per-dataset breakdown is on the benchmarks page.
What that number doesn't capture is that lecture errors cluster in exactly the words you care about most: the technical vocabulary, the researcher names, the course-specific jargon. A general-purpose model has heard "protein" a million times and "phosphofructokinase" rather less. Everyday words come through cleanly and the terms you'd actually want to search for are where mistakes land.
Which is a solvable problem, and the fix takes about two minutes.
Prime the model with your syllabus
Universal-3.5 Pro supports contextual prompting — you give the model the domain and vocabulary it's about to encounter, before it encounters it.
For a lecture, the source material already exists. Paste in the course title and description, the week's reading list, the section headings from the syllabus, the lecturer's name and the names of researchers whose work the course covers. If there's a glossary in the textbook, that's the single highest-value thing you can supply.
This isn't a generic instruction to "be accurate." It's giving the model the prior that a well-prepared human transcriptionist would have from having read the syllabus. In healthcare, feeding the model a patient's prior-visit note cut missed medical terms by 31% in our internal testing — the mechanism for a chemistry lecture and a reading list is the same one.
Set it up once per course and reuse the same prompt for every lecture in the semester. The API documentation covers the parameter shape.
Run the same lecture twice — once plain, once primed with your syllabus glossary — and compare how the technical terms come through.
Lectures in other languages, and for international students
Universal-3.5 Pro handles 18 languages out of the box with native code-switching — English, Spanish, French, German, Italian, Portuguese, Arabic, Danish, Dutch, Finnish, Hebrew, Hindi, Japanese, Mandarin, Norwegian, Swedish, Turkish, and Vietnamese. Code-switching matters more in a lecture hall than people expect. A visiting lecturer drops into their first language for a technical term, a seminar in Barcelona moves between Spanish and English mid-sentence, and a model that has to be told which language it's hearing will mangle both.
Beyond those 18, Universal-2 covers 99+ languages at $0.15/hr. And for pre-recorded audio, translation is available into 100+ target languages, so a lecture recorded in German can be handed to a student who reads Mandarin.
Education platforms operating at that scale build on this directly — Varsity Tutors, part of the publicly listed Nerdy, and DeepLearning.AI, Andrew Ng's platform. You can see more about the teams building on the API.
What audio event tagging adds to a lecture recording
A lecture recording contains more than speech. Speech understanding features include audio event tagging across 100+ event types — applause, laughter, music, silence, background noise — which turns out to be a useful index for a long recording.
A burst of laughter usually marks an aside. A long silence usually marks the professor writing something on the board that isn't in the transcript at all, which is precisely the moment to check your photo of the whiteboard. Note that audio event tags are now removed from transcripts by default, so you'll need to opt in if you want them inline.
Pair that with speaker diarization and a seminar transcript stops being a wall of text — you can see who said what, and skip straight to the student questions.
Are you allowed to record a lecture?
Usually yes, often with conditions, and the institution's policy is stricter than the law.
Recording consent rules vary by US state, and some require consent from everyone being recorded rather than just one party. But that's rarely the binding constraint on campus, because most universities have their own policy — typically permitting recording for personal study, prohibiting redistribution, and sometimes requiring you to delete recordings at the end of term. Check your student handbook, and if you're recording under a disability accommodation, the accommodation letter usually settles the question. Ask the lecturer anyway. It takes ten seconds and it's the difference between a courtesy and a complaint.
Recording setup, briefly
Your phone is fine. Sit in the front third of the room, put the device on the desk with nothing on top of it and the microphone pointing toward the lecturer, and check the battery before a three-hour lab. A cheap clip-on lavalier improves things noticeably if you can get it near the front. None of this needs a budget — room position beats equipment almost every time.
What a semester of lectures actually costs
Less than a textbook:
| Scenario | Audio hours | Approx. cost |
|---|---|---|
| One 50-minute lecture | ~0.8 | ~$0.18 |
| Four courses, 15-week semester | 180 | ~$38 |
| Institutional back catalogue | 1,000 | ~$210 |
For an institution captioning a back catalogue, the same arithmetic scales predictably — a thousand hours is around $210 at list price. Live captioning is priced separately at the streaming rate; the full pricing breakdown has the details.
The accessibility case is the whole case
Here's what's easy to miss. Every argument for lecture transcription as an accommodation — searchable content, no penalty for a lapse in attention, the ability to review at your own speed — describes something every student benefits from. The accommodation is just the version where the need is documented.
Which is a reasonable way to decide how to build. If your captions work for the student who can't hear the lecture, they'll work for the student who missed eight minutes of it, the one who's studying in their third language, and the one revisiting week three in April. Build for the hardest case and the rest comes free.
Whether it's one course or a decade of archived recordings, start with an API key and transcribe your first hour today.
Frequently asked questions
How accurate is lecture transcription when the lecturer has a strong accent?
Modern speech recognition models are trained on a wide range of speakers and handle most accented speech well, though performance does vary by accent and audio quality. The bigger lever is usually contextual prompting: giving the model the course vocabulary in advance fixes far more errors than accent alone accounts for. Test a single lecture before committing to a semester.
How do I get lecture transcription as a disability accommodation?
Register with your institution's disability services office and provide documentation from a qualified professional, then request lecture transcription or captioning specifically rather than a general accommodation. Put the request in writing so there's a record of what was asked and what was granted. If the accommodation you receive turns out to be insufficient in practice, you can request a revision — this is expected and routine.
Is recording lectures allowed?
In most cases yes, but your university's policy governs, and it's typically stricter than state recording law. The common pattern permits recording for personal study while prohibiting redistribution, and some institutions require deletion at the end of term. If you're recording under a documented accommodation, your accommodation letter usually authorizes it explicitly.
Can I transcribe lectures delivered in languages other than English?
Yes. Universal-3.5 Pro covers 18 languages with native code-switching, which handles the common case of a lecturer moving between languages mid-sentence. For broader coverage, Universal-2 supports 99+ languages, and pre-recorded audio can be translated into 100+ target languages after transcription.
Should I use live captioning or transcribe the lecture afterwards?
Use both if accessibility is the goal, because they solve different problems. Live captioning makes the lecture followable in the room as it happens, while post-lecture transcription produces the searchable artifact you study from and can be run on the highest-quality recording available. Live captioning is billed on session duration; post-lecture transcription is billed on audio hours.
What does it cost to transcribe a full semester of lectures?
At $0.21 per audio hour, a typical four-course semester of about 180 lecture hours costs roughly $38. Billing is per second with no minimums, so short lectures and cancelled classes cost proportionally less. Institutions captioning archived recordings can estimate at the same rate — around $210 per thousand hours at list price.
Lorem ipsum dolor sit amet, consectetur adipiscing elit, sed do eiusmod tempor incididunt ut labore et dolore magna aliqua. Ut enim ad minim veniam, quis nostrud exercitation ullamco laboris nisi ut aliquip ex ea commodo consequat. Duis aute irure dolor in reprehenderit in voluptate velit esse cillum dolore eu fugiat nulla pariatur.




