MP3, M4A (the format iPhone voice memos use), WAV, WebM, and OGG upload directly. Video files like MP4 and MOV work too — Scholarly converts the audio track to text and ignores the picture. Long recordings are fine, including multi-hour lectures, and you don't need to split them first.
Converting audio to text gives you the verbatim record — every sentence, in order, with timestamps and speaker turns. That's what you want when the exact wording matters: a quote for an essay, the precise phrasing of a definition, or who said what in a seminar. If you'd rather have the recording condensed and organized by topic instead, use Audio to Notes — it runs the same transcription underneath, then summarizes it into structured study notes. Either way the full transcript stays attached, so you never lose the source text.
On clear audio, accuracy is high — clean lecture and interview recordings come back close to word-for-word. The transcript is timestamped, so when a single word looks off you can click straight to that second and confirm it against the audio in a couple of taps, rather than re-reading the whole thing.
What genuinely hurts accuracy is the recording, not the model: a microphone far from the speaker, heavy crosstalk, thick accents on top of background noise, or lots of specialized jargon all degrade the text. For anything high-stakes — a number, a name, a technical term you'll quote — verify it against the audio at its timestamp. It takes seconds, and you're checking against what was actually said.
The recording becomes a source in your Scholarly workspace, so everything builds on the transcript without re-uploading: search the full text for any term, turn it into clean study notes, generate spaced-repetition flashcards or a practice quiz, or ask the AI chat questions and get answers that cite the exact passage of the audio. Text is more useful than audio precisely because you can act on it — and reading a transcript while you study reinforces understanding far better than passively replaying a recording.
Need the verbatim text first? Our lecture transcription tool is built for exactly that — a searchable, timestamped transcript of a full class, with speaker turns kept intact.