- Home
- How-to Guides
- Anki Decks With Audio
How to Make an Anki Deck With Audio
A vocabulary card without sound teaches you a word you cannot say and would not recognise if someone said it to you. There are four ways to fix that in Anki, they differ more than they look, and one of them requires nothing to be installed at all.
Four ways to get audio into Anki
A deck that already contains the audio
- This is what NextLang generates: pick your language pair, generate the deck, and the pronunciation of the learned-language side is packed into the same package.
- Neural voices across all 20 supported languages, one voice per language, chosen by blind comparison rather than by price.
- The audio survives sync, works offline, and works identically on AnkiDroid and AnkiMobile, because it is just media in your collection.
- Nothing to maintain: no API key in an add-on, no service that can start charging or shut down between one deck and the next.
An add-on that generates audio inside Anki
- Powerful and configurable: many providers, many voices, bulk generation over an existing deck.
- Desktop only. The add-on runs in Anki for computers, so audio has to be generated there and synced to your phone afterwards.
- The good voices need an account with a cloud provider and a key pasted into the add-on. The free voices are noticeably worse, which is the whole reason this category exists.
- Worth it when you already own a large deck without audio. Redundant when the deck you are about to make can arrive with audio in it.
Anki's own text-to-speech tag
- No files, so the collection stays small and any change to the text is spoken immediately.
- The voice is whatever the operating system provides, which varies by platform and is usually a step below a neural cloud voice.
- It is a template edit, so it applies to a note type rather than to a card, and a missing voice on one device means silence on that device only.
- A reasonable fallback for a deck you built yourself, and the only route that costs nothing at all.
Buying a pre-made audio deck
- You get a finished product, vetted, with a word list someone thought about.
- You do not get your material: the words from your course, your book, your show. A bought deck teaches its own selection.
- Fine as a base layer. It does not remove the need for a deck built from what you are actually reading.
The four routes side by side
All four can add audio to a whole deck automatically rather than card by card. What separates them is where the sound lives afterwards, and what it costs to keep it.
| Route | Automatic in bulk | On phone | Offline | Voice | Cost |
|---|---|---|---|---|---|
| Audio already inside the .apkg | Yes, the whole deck at once | Yes | Yes | Neural, one per language | Free, no credit |
| Add-on (HyperTTS, AwesomeTTS) | Yes, bulk over a selection | Files sync, the add-on does not run there | Yes, once generated | Your choice, best ones need a provider key | Add-on free or paid, provider billed separately |
| Anki's own {{tts}} tag | Yes, it is a template rule | Yes, if the device has the voice | Depends on the device voice | Whatever the operating system ships | Free |
| A bought audio deck | Nothing to do, it arrives finished | Yes | Yes | Neural or native speaker | A few dollars per deck |
The row that decides it for most people is the phone. An add-on runs inside Anki for computers, so the audio has to be made there first and synced afterwards; a package that already carries its mp3 files plays everywhere from the first review.
Adding the audio automatically, step by step
Four steps, no add-on and no per-card work: the whole deck is voiced in one pass, and the only decision that matters is whether to leave the audio toggle on.
Open the Anki generator, choose the language you are learning and the one you want translations in, set your CEFR level and the card count. Word pairs or cloze sentences both support audio.
The deck is shown before anything is downloaded, and any card can be edited. This is also where you decide whether the examples are worth keeping.
The toggle sits in the download menu and is remembered between exports. On, the file carries the mp3 files and weighs a few hundred kilobytes; off, it is a couple of kilobytes of text.
Double-click the file, or File then Import. The deck arrives with the note type, the tags and the media already in place, and the speaker plays on the answer side from the first review.
The same audio travels in a native .mochi package if you study in Mochi, and the words you saved in your vocabulary export the same way - including the ones imported from a subtitle file.
When the sound does not play
Audio problems in Anki are nearly always about where the file is, not about whether the card is right.
That is Anki telling you it has the reference but not the file. It happens when a deck is shared as text rather than as a package: a CSV cannot carry media. Re-export as .apkg, or check that the import was the package and not a text file.
The media has not synced yet. Media syncs separately from cards and can lag behind, especially on a first sync of a large collection. Force a sync on both ends and give it a minute.
Audio is the reason: a text deck is kilobytes, an audio deck is hundreds. If you only wanted the text, turn Include audio off before downloading - the setting is remembered for the next export too.
The audio belongs on the side in the language you are learning, and that is where it is generated. If your template shows that side first, you are hearing the answer before the question - swap the template or move the tag, rather than regenerating the deck.
Word audio is included for every account. Example-sentence audio is a premium extra, and cloze decks are the exception that proves the rule: a cloze card is a sentence, so the sentence is what gets voiced.
One voice per language, chosen for clarity rather than for a region, and the region follows from the voice: Spanish is Latin American and Portuguese is Brazilian. Picking a voice per deck is not available yet; if the accent matters to you more than the convenience does, an add-on with a provider account is the honest answer.
What embedding does to your collection
An embedded file behaves differently from a generated one, and the difference shows up months later - in sync, in backups, and on the day a service you depended on changes its terms.
The mp3 files land in your collection.media folder and are referenced from the field by name. That means they are yours: they sync, they back up, they survive the deck being renamed or reorganised, and they keep working if this site disappears. It also means duplicates cost you nothing on our side and a little on yours - the same word in two decks is one file for us, two references for Anki, and identical bytes either way.
Every recording is content-addressed by a hash of the text and the voice, so the first learner to need a word pays for the synthesis and everyone after that gets the same file. At the scale of a shared vocabulary, that means near-zero marginal cost, which is why audio is not metered here while the rest of the market sells it by the sentence.
Once the audio is in the package, turning the deck into listening practice is a template change rather than a new deck: move the audio reference to the question side and hide the text, and you have a card that plays a word and asks you to recognise it. Do it on a copy of the note type if you want both kinds of review from the same notes.
On a cloze deck the recording is of the whole sentence with the markers stripped, so you hear a natural line rather than a sentence with a hole in it. Keeping it on the answer side gives you the confirmation you want after producing the word; moving it to the question turns the card into dictation, which is harder and better if you can take it.
The .mochi package carries its media the same way, so a generated Mochi deck speaks without a Mochi Pro subscription and without any add-on, which no other generator currently does. If you study on both apps, generate the deck twice rather than converting the file.
On cloze decks the recording is the whole sentence, which changes what the card can do - the cloze guide covers that combination.
Premium access includes:
Frequently asked questions
Not for a deck generated here: the mp3 files are already inside the package. An add-on is for adding audio to a deck that does not have any, which is a different problem.
Yes. The files live in your collection, so reviews work on a plane or on the underground exactly like any other card.
No. Audio never consumes a credit, on any plan. One credit is charged for the generation itself, whatever the card count and whether or not you take the audio.
All 20 languages the generator supports, with a neural voice for each. The voice speaks the side in the language you are learning.
Not through NextLang: the audio actions only voice cards from your own results and your own vocabulary, by design. For an existing collection, an add-on inside Anki is the right tool.
Keep reading
Related guides that build on what you just read.