MP3, M4A (the format iPhone voice memos use), WAV, WebM, and OGG upload directly. Video files like MP4 and MOV work too — Scholarly transcribes the audio track and ignores the picture. Long files are fine, including multi-hour seminar recordings.
Transcription gives you the verbatim record — every sentence, in order. Notes are the condensed version: organized by topic, with definitions and takeaways pulled out. If what you mainly need is the verbatim text with timestamps and speaker turns, use Lecture Transcription instead. Audio to Notes runs the same transcription underneath, then does the organizing for you — and keeps the transcript attached either way, so you never lose the source.
Better than you'd expect, with honest limits. Spoken explanations are naturally repetitive and out of order — that's exactly what the note-generation step fixes, because it organizes by topic rather than by the order things were said. A rambling 40-minute memo usually becomes a tight one-page outline.
What genuinely hurts quality is bad audio, not bad structure: a microphone far from the speaker, heavy crosstalk, or loud background noise degrade the transcript the notes are built on. For anything high-stakes, check critical numbers and terms against the timestamped transcript — it takes seconds, and you're verifying against what was actually said.
The recording becomes a source in your Scholarly workspace, so everything builds on it without re-uploading: generate spaced-repetition flashcards or a practice quiz from the notes, ask AI Meeting Notes questions and get answers that cite the exact passage of the recording, or combine it with your PDFs and slides into one study set for the exam.
New to the workflow? Our guide on how to turn lecture recordings into study notes walks through every step from capturing the audio to building flashcards and a practice exam.