The average professor speaks at somewhere between 120 and 150 words per minute.1 The average person types at around 40 words per minute.2 Nobody talks about this gap like it's a problem, but it is — it's the reason your notes from a two-hour lecture are either shallow bullet points or an illegible sprint-typing disaster that you never actually review.
Copying slides doesn't fix it. Slides give you the structure of what was said, not the content. The explanations, the worked examples, the "here's why this actually matters" tangents — all of that lives in the audio, not the deck. You can walk out of a lecture with perfect slide notes and still have no idea what was being taught.
There's a real solution to this, and it doesn't require a $400 recording device or a complicated setup. This post walks through a workflow that works for most lecture environments and doesn't cost much to run.
What People Try First (and Why It Partially Works)
Before getting to the workflow, it's worth being honest about the tools students already reach for.
Cloud-based transcription apps are the most common answer you'll find in Reddit threads and "best apps for college students" listicles. They work reasonably well for transcription accuracy. The problem is pricing: free tiers cap at a few hundred minutes per month, which disappears fast if you're in four or five courses with weekly lectures. Paid tiers run $15–17/month. That's not nothing on a student budget, and everything uploads to remote servers — your lecture audio, someone else's intellectual property, potentially your own research discussions.
Google Docs voice typing is free and surprisingly accurate for a browser tool. The dealbreaker is that it requires an active microphone input and a browser tab. You can't record a lecture passively while taking notes by hand, you can't run it on your phone in your pocket, and you can't process audio you recorded earlier. It's designed for live dictation at a desk, not for capturing a 90-minute environmental biology lecture.
OpenAI Whisper (the open-source model) is technically more accurate than most commercial tools. If you have the patience to set it up, run it on a machine that can handle it, and wait for processing to complete, it produces solid output. The realistic version of this workflow for most students involves downloading a model, installing Python dependencies, and waiting 20 minutes for a transcription to finish on a laptop that really isn't built for it. Doable. Not practical.
The Two-Step Verity Workflow
The approach that works is simpler: record on your phone during the lecture, process on your laptop later.
Step one: record with the Verity app. The iOS and Android apps let you start a recording session and leave it running. Your phone sits in your bag or on your desk, picking up the lecture. You're free to take handwritten notes, follow along on slides, or just listen. You're not trying to type everything at once.
Step two: open Verity on your laptop, load the recording, and run it through the formal persona. Verity transcribes and polishes on its own servers — audio is processed in memory and immediately discarded. Nothing is stored. The result is not a raw transcript. It's a processed version that reads like someone wrote it down.
The cross-device sync means your recordings move from phone to laptop without manual file transfer. Start the recording, end the class, let Verity process in the background while you grab coffee.
Try Verity free
Record on your phone. Study on your laptop. Free during launch — all features included.
Download FreeWhy the AI Polish Step Matters for Academic Notes
Raw transcription of a lecture is not the same as a useful transcript of a lecture.
Spoken academic language is full of verbal scaffolding. Phrases like "so what I mean by that is," "going back to what I said earlier," and "and this is the key point here" don't carry information — they're navigation signals that work in audio and fall completely flat in text. A raw transcript preserves all of them, which means you're reading twice as many words to extract the same amount of content.
There's also the issue of informal grammar. Professors talk the way humans talk: sentence fragments, back-tracking, mid-thought corrections, dependent clauses that never close. This is fine to listen to. It's confusing to read verbatim.
Verity's formal persona is specifically designed to restructure spoken language into clean written prose. It removes the filler, resolves the grammar, and produces sentences that read like they were meant to be read. The conceptual content stays intact. The verbal scaffolding disappears.
For studying, this matters. Notes you can actually read are notes you'll actually use. A 90-minute lecture that produces three pages of clean, structured prose is more useful than eight pages of half-formed sentences and "um, so anyway."
Other Ways Students Use It
Lecture transcription is the obvious use case, but it's not the only one.
Voice memos to essay drafts. If you're the kind of person who thinks better out loud, Verity can handle the gap between talking through an argument and having a written version of it. Record yourself explaining the thesis of your paper while walking between classes. Run it through the formal persona. You'll usually end up with a rough draft that's more coherent than anything you'd type staring at a blank document.
Converting office hours into notes. Office hours conversations are some of the most information-dense academic interactions most students have. They're also almost never documented. Asking a professor to repeat themselves five times while you type is awkward. Recording and transcribing later is not.
Review sessions. Study groups often produce better explanations of material than lectures do — students have usually figured out which parts are confusing and how to explain them. Recording a group review session and processing it through Verity gives you a set of notes in the language your peers actually use to explain the material, which is often more useful for exam prep than the formal lecture version.
On Privacy and Lecture Content
Lecture audio is not a neutral recording. It often contains a professor's original explanations, examples they've developed over years, and sometimes unpublished research discussion. The question of who owns that content — and where it lives after you record it — is worth thinking about.
Cloud-based transcription tools upload your audio to external servers. This creates real questions: Is the audio stored? For how long? Under what terms? Could it be used to train models? Most terms of service are vague enough that the honest answer is "unclear."
Verity processes audio on its own servers and immediately discards it. Nothing is stored, nothing is shared, and no third-party AI providers are involved. For lecture content — where a professor hasn't consented to cloud storage — zero retention is the cleanest answer.
The Free Tier Is Actually Useful Here
Verity is free during launch — all features included, no word limits, no credit card.
A typical 50-minute lecture produces roughly 6,000 to 7,000 words of spoken content.5 After AI polish, the cleaned output is usually significantly shorter — verbal scaffolding removal alone can cut 20 to 30 percent. Realistically, a processed lecture might produce 3,000 to 4,500 words of usable notes.
During launch, there are no limits on how many lectures you process. Transcribe every class, every week. When pricing kicks in after launch, a free tier will still be available.
Practical Tips for Better Results
A few things that make a meaningful difference in transcription quality:
Positioning matters more than equipment. You don't need an external microphone. You do need your phone close enough to the speaker to get clear audio. Sitting in the first few rows and placing your phone on the desk (face down to reduce notification interference) works well. The back of a lecture hall with an overhead PA system can be tricky — the audio often sounds fine in person but compresses badly in recordings.
Use review sessions differently. When you're working through your own notes for exam prep, switch from recording to live dictation. Talk through what you're reviewing — explain concepts in your own words, flag what you're uncertain about, summarize what you've understood. This creates a second set of notes that's organized around your own comprehension rather than the lecture structure.
Batch processing works. Don't feel like you need to transcribe each lecture immediately. Letting recordings accumulate across the week and processing them all on a Sunday evening is a perfectly valid workflow. Verity handles batch processing without needing you to babysit each file.
Match the persona to the content. The formal persona is right for most lecture material. For quick voice memos — reminders, ideas, rough outlines — the Minimal persona is faster and doesn't restructure as aggressively. Use formal when you need the output to actually read well.
The note-taking problem in lectures is structural — you can't type at the speed of speech, and you shouldn't have to. Recording and processing is how you close that gap without sacrificing comprehension for speed. The workflow described here isn't complex, and for most students it takes about a week of habit formation before it becomes automatic. (For general tips on getting faster with dictation, see our beginner's speed guide.)
If you want to try it, the free tier is a reasonable way to start without any commitment. One lecture. See if the output is actually useful for studying. If it is, you'll know.
Try Verity free
Record on your phone. Study on your laptop. Free during launch — all features included.
Download Free- Baruch College — Speaking Rate Guide — ASHA benchmark: ~130 WPM for formal speech. ↵
- Aalto University (CHI 2018) — Average typing speed of 52 WPM across 168,000 participants. ↵
- Derived: 130 WPM (ASHA formal speech rate) × 50 minutes = 6,500 words. ↵