- Home
- How-to Guides
- Subtitles to an Anki Deck
How to Turn Subtitles Into an Anki Deck
A subtitle file is the best vocabulary source you already own: it is real dialogue, at natural speed, from something you chose to watch. This guide takes an .srt or .vtt from your disk to a finished Anki package with the pronunciation audio inside it, and is honest about the one step that stays manual.
From subtitle file to deck, step by step
Get the subtitle file
- On YouTube: open the video, click the three dots under it, choose Show transcript, then select the text and copy it. The panel is part of the player, so nothing is scraped and nothing breaks.
- From a course, a podcast or a film you own: the subtitle track usually ships as a separate file already. Any of them works.
- Supported formats: .srt, .vtt, .sbv, .ass, .ssa. Plain .txt and .md work too, which is what a copied transcript becomes when you paste it.
- Pick the track in the language you are learning, not the translated one. A Spanish film with English subtitles teaches you English word frequencies.
Open the dictionary for your language pair
- Go to My Vocabulary, pick the pair (or create it by saving a first word), then choose Import from text.
- The source language is the language spoken in the video. The target language is the one you want the translations in.
- Set your CEFR level, A1 to C1. This is the setting that decides whether the import returns 'to run' or 'to make the most of'.
Upload the file
- One import reads up to 60,000 characters, which covers a feature film or roughly two hours of talking. A whole season goes in one episode at a time.
- Ask for up to 50 candidate words on free credits, or up to 150 once you have bought credits. One credit per import, refunded if nothing usable comes back.
- The words come back as lemmas, not as they appeared on screen: 'was running' is returned as 'to run', which is the form worth learning.
- Each candidate carries a sentence taken from the subtitles themselves. That sentence is a line from the video, not something the model invented.
Keep the words worth keeping
- Words already in this dictionary are skipped automatically, so importing a second episode of the same show adds only what is new.
- Names, places and brand names are vocabulary in the technical sense and useless in practice. Drop them.
- Keep the words you almost knew. A word you met once and could not place is the one spaced repetition pays off on; a word you have never seen and cannot use in a sentence yet usually is not.
- Nothing is saved until you press save, and nothing costs a credit at this step.
Export the deck
- Search and the Learning / Learned filters decide what goes into the file, so you can export only this video's words, or everything you have saved for the pair.
- Anki package (.apkg) carries the note type, the fields and the tags with it, so there is no import dialog, no separator to choose and no columns to map.
- Pronunciation audio is embedded in the package itself, so the cards keep speaking inside Anki with no text-to-speech add-on installed.
- The export costs no credit. File exports are unlocked by having bought credits at some point; importing, curating and studying on the site are free.
Import into Anki and study
- FSRS treats a generated deck exactly like a hand-made one. No special export settings are needed.
- The audio plays on the answer side by default. If you want a listening card instead, change the card template once and every note follows.
- On phones: AnkiDroid and AnkiMobile both open .apkg files directly, or sync from the desktop as usual.
What goes wrong with subtitles
Subtitle files are messy in predictable ways. Most of the mess is handled for you; the rest is a decision only you can make.
YouTube's automatic captions arrive as one long unpunctuated stream, so example sentences come out longer and blunter than they should. The word list itself survives this fine - lemmas do not depend on commas. If a channel offers a human-written track, take that one instead.
That is what a .srt looks like from the inside, and it is handled: cue numbers, timecodes and the blank lines between cues are stripped before the text is read. You do not need to clean the file up first.
A single import reads 60,000 characters. A film fits; a full season does not. Split it by episode, which is the better unit anyway: the words from one episode arrive with the sentences from that episode.
Subtitles are dialogue, and dialogue is full of names. The extractor drops obvious proper nouns, but a name that doubles as a common word slips through. This is a thirty-second fix during curation and the main reason curation is not automated away.
That is the CEFR level, not the video. The import returns words at the level you set and one level above it, and leaves out anything a learner at that level already knows. If a B2 import feels trivial, raise the level rather than the word count.
A dubbed film often ships with subtitles that were translated separately, so they do not match the spoken track word for word. For vocabulary that is harmless. For sentence mining, where you want the sentence you actually heard, pick the track that belongs to the original audio.
Why there is no "paste a YouTube link" field
Plenty of flashcard tools advertise a box where you drop a video link. It is a genuinely nicer first screen, and we decided against it on purpose.
Fetching a transcript from a server is no longer a solved problem. YouTube now expects a proof-of-origin token and blocks requests from the data-centre addresses every hosted app runs on, so a link field works on a developer's laptop and fails the day it is deployed. The tools that still offer one either pay a third-party transcript API per request or route traffic through residential proxies. Both put a brittle dependency between you and your deck, and both quietly break.
Copying the transcript out of the panel under the video takes about fifteen seconds and cannot break. That is why step 1 is the way it is: the manual step buys a route that keeps working, and everything after it is automated.
The tools in the language-learning space that handle video properly - Language Reactor, Migaku, subs2srs - are browser extensions or desktop software for exactly this reason: the transcript is fetched by your own browser, on the page, in your own session. If we ever ship an extension, the link route arrives with it, and this guide gains a shortcut rather than a rewrite.
In the meantime the rest of the pipeline is unchanged: the same Anki generator writes cards from a topic when you have no file at all, and the free CSV to Anki converter handles a word list you already keep in a spreadsheet.
Sentence mining: i+1, cloze and deck hygiene
Mining vocabulary from what you watch is an old practice with a small number of rules that actually matter. These are the ones that change how you use the import.
A word learned in isolation gives you a translation. A word learned inside a line you have heard gives you the collocation, the register and a hook to recall it by. This is the core of sentence mining as Migaku, subs2srs and the Yomitan crowd practise it, and it is why the example sentence is pulled from your subtitle file rather than generated: you have heard that sentence, in that voice, in that scene.
Sentence mining has one durable rule: mine sentences where exactly one element is unknown. Everything else is comprehensible, so the new word is anchored by context you already own. Setting your CEFR level does this at the word level - the import returns items at your level and one step above, and drops what you already know - which is why an honest level setting matters more than a high card count.
A word card asks 'what does this mean'. A cloze card asks you to produce the word in a sentence with the answer hidden, which is closer to what speaking demands. The Anki generator writes native Cloze note type cards with one blank per sentence, so the count stays predictable and each note makes exactly one card. Use the sentences from an episode as the source text and you have a cloze deck built from dialogue you have already watched.
Import episode by episode into the same dictionary and the duplicate check does the deck maintenance for you: episode 4 contributes only words episodes 1 to 3 did not. Export once a week rather than once per episode, and you get one deck that grows instead of eleven small decks with overlapping cards - which is the failure mode that makes people abandon Anki.
The mp3 files inside the package are neural text-to-speech recordings of the word in the language you are learning, not the actor's voice from the video, which we neither have nor have the right to ship. Word audio is free for every signed-in account; audio for example sentences is a premium extra. If you want the original audio, that is subs2srs territory and it needs the media file on your own machine.
Premium access includes:
Frequently asked questions
No, and that is deliberate rather than a gap. Fetching a transcript from a server now requires either a paid third-party API or residential proxies, both of which put a fragile dependency between you and your deck. Copying the transcript from the panel under the video takes about fifteen seconds and never breaks.
.srt, .vtt, .sbv, .ass, .ssa, plus plain .txt and .md for a transcript you copied out of a page. All of them are cleaned of timecodes and cue markers before the text is read.
One credit per import, whatever the file size, and it comes back if nothing usable is found. Choosing which words to keep is free, studying them on the site is free, and the deck export itself costs nothing - file exports are unlocked by having bought credits at any point.
Yes. The mp3 files travel inside the .apkg, so the deck plays in Anki with no add-on, no AwesomeTTS and no internet connection while you study. The same words also export to Mochi as a native .mochi package with the audio embedded, which does not require Mochi Pro.
Yes. The import page takes a file from your phone's file picker, and NextLang installs as an app on Android and iOS. Getting the subtitle file is the fiddly part on mobile; the rest works the same as on desktop.
Keep reading
Related guides that build on what you just read.