How to Make an Anki Deck With Audio

A vocabulary card without sound teaches you a word you cannot say and would not recognise if someone said it to you. There are four ways to fix that in Anki, they differ more than they look, and one of them requires nothing to be installed at all.

By Siarhei Hamanovich
The short version
Generate the deck here and the mp3 files are packed inside the .apkg. Open the file in Anki and the cards speak - on desktop, on Android, on iOS, offline, with no add-on, no API key and no credit spent on the audio.

Four ways to get audio into Anki

A deck that already contains the audio

The mp3 files travel inside the .apkg. You open the file and the cards speak, on every device, with nothing installed and nothing configured.
  • This is what NextLang generates: pick your language pair, generate the deck, and the pronunciation of the learned-language side is packed into the same package.
  • Neural voices across all 20 supported languages, one voice per language, chosen by blind comparison rather than by price.
  • The audio survives sync, works offline, and works identically on AnkiDroid and AnkiMobile, because it is just media in your collection.
  • Nothing to maintain: no API key in an add-on, no service that can start charging or shut down between one deck and the next.

An add-on that generates audio inside Anki

AwesomeTTS, its successor HyperTTS, and similar add-ons add a button to the editor that fills a field with generated speech.
  • Powerful and configurable: many providers, many voices, bulk generation over an existing deck.
  • Desktop only. The add-on runs in Anki for computers, so audio has to be generated there and synced to your phone afterwards.
  • The good voices need an account with a cloud provider and a key pasted into the add-on. The free voices are noticeably worse, which is the whole reason this category exists.
  • Worth it when you already own a large deck without audio. Redundant when the deck you are about to make can arrive with audio in it.

Anki's own text-to-speech tag

Anki can speak a field at review time with a {{tts}} tag in the card template, using the voices installed on the device.
  • No files, so the collection stays small and any change to the text is spoken immediately.
  • The voice is whatever the operating system provides, which varies by platform and is usually a step below a neural cloud voice.
  • It is a template edit, so it applies to a note type rather than to a card, and a missing voice on one device means silence on that device only.
  • A reasonable fallback for a deck you built yourself, and the only route that costs nothing at all.

Buying a pre-made audio deck

Several shops sell curated language decks with native or neural audio baked in, usually a few dollars per deck.
  • You get a finished product, vetted, with a word list someone thought about.
  • You do not get your material: the words from your course, your book, your show. A bought deck teaches its own selection.
  • Fine as a base layer. It does not remove the need for a deck built from what you are actually reading.

The four routes side by side

All four can add audio to a whole deck automatically rather than card by card. What separates them is where the sound lives afterwards, and what it costs to keep it.

Ways to add audio to Anki cards compared by automation, mobile support, offline playback, voice quality and cost
RouteAutomatic in bulkOn phoneOfflineVoiceCost
Audio already inside the .apkgYes, the whole deck at onceYesYesNeural, one per languageFree, no credit
Add-on (HyperTTS, AwesomeTTS)Yes, bulk over a selectionFiles sync, the add-on does not run thereYes, once generatedYour choice, best ones need a provider keyAdd-on free or paid, provider billed separately
Anki's own {{tts}} tagYes, it is a template ruleYes, if the device has the voiceDepends on the device voiceWhatever the operating system shipsFree
A bought audio deckNothing to do, it arrives finishedYesYesNeural or native speakerA few dollars per deck

The row that decides it for most people is the phone. An add-on runs inside Anki for computers, so the audio has to be made there first and synced afterwards; a package that already carries its mp3 files plays everywhere from the first review.

Adding the audio automatically, step by step

Four steps, no add-on and no per-card work: the whole deck is voiced in one pass, and the only decision that matters is whether to leave the audio toggle on.

1. Generate the deck

Open the Anki generator, choose the language you are learning and the one you want translations in, set your CEFR level and the card count. Word pairs or cloze sentences both support audio.

2. Check the cards on screen

The deck is shown before anything is downloaded, and any card can be edited. This is also where you decide whether the examples are worth keeping.

3. Leave Include audio on

The toggle sits in the download menu and is remembered between exports. On, the file carries the mp3 files and weighs a few hundred kilobytes; off, it is a couple of kilobytes of text.

4. Download the .apkg and open it

Double-click the file, or File then Import. The deck arrives with the note type, the tags and the media already in place, and the speaker plays on the answer side from the first review.

The same audio travels in a native .mochi package if you study in Mochi, and the words you saved in your vocabulary export the same way - including the ones imported from a subtitle file.

When the sound does not play

Audio problems in Anki are nearly always about where the file is, not about whether the card is right.

The card shows [sound:...] as text

That is Anki telling you it has the reference but not the file. It happens when a deck is shared as text rather than as a package: a CSV cannot carry media. Re-export as .apkg, or check that the import was the package and not a text file.

Audio plays on the desktop but not on the phone

The media has not synced yet. Media syncs separately from cards and can lag behind, especially on a first sync of a large collection. Force a sync on both ends and give it a minute.

The deck is much bigger than expected

Audio is the reason: a text deck is kilobytes, an audio deck is hundreds. If you only wanted the text, turn Include audio off before downloading - the setting is remembered for the next export too.

The wrong side is being spoken

The audio belongs on the side in the language you are learning, and that is where it is generated. If your template shows that side first, you are hearing the answer before the question - swap the template or move the tag, rather than regenerating the deck.

I want the sentence spoken, not just the word

Word audio is included for every account. Example-sentence audio is a premium extra, and cloze decks are the exception that proves the rule: a cloze card is a sentence, so the sentence is what gets voiced.

The audio does not match the accent I am learning

One voice per language, chosen for clarity rather than for a region, and the region follows from the voice: Spanish is Latin American and Portuguese is Brazilian. Picking a voice per deck is not available yet; if the accent matters to you more than the convenience does, an add-on with a provider account is the honest answer.

What embedding does to your collection

An embedded file behaves differently from a generated one, and the difference shows up months later - in sync, in backups, and on the day a service you depended on changes its terms.

What embedding actually does to your collection

The mp3 files land in your collection.media folder and are referenced from the field by name. That means they are yours: they sync, they back up, they survive the deck being renamed or reorganised, and they keep working if this site disappears. It also means duplicates cost you nothing on our side and a little on yours - the same word in two decks is one file for us, two references for Anki, and identical bytes either way.

Why the audio costs no credit

Every recording is content-addressed by a hash of the text and the voice, so the first learner to need a word pays for the synthesis and everyone after that gets the same file. At the scale of a shared vocabulary, that means near-zero marginal cost, which is why audio is not metered here while the rest of the market sells it by the sentence.

Listening cards from the same notes

Once the audio is in the package, turning the deck into listening practice is a template change rather than a new deck: move the audio reference to the question side and hide the text, and you have a card that plays a word and asks you to recognise it. Do it on a copy of the note type if you want both kinds of review from the same notes.

Audio and cloze together

On a cloze deck the recording is of the whole sentence with the markers stripped, so you hear a natural line rather than a sentence with a hole in it. Keeping it on the answer side gives you the confirmation you want after producing the word; moving it to the question turns the card into dictation, which is harder and better if you can take it.

Mochi gets the same treatment

The .mochi package carries its media the same way, so a generated Mochi deck speaks without a Mochi Pro subscription and without any add-on, which no other generator currently does. If you study on both apps, generate the deck twice rather than converting the file.

On cloze decks the recording is the whole sentence, which changes what the card can do - the cloze guide covers that combination.

Premium Content
Unlock the complete guide with all advanced techniques

Premium access includes:

Complete guides with all sections unlocked
Import-ready files: .apkg and .mochi packages, plus CSV, TSV, TXT and MD
Bigger sets: up to 30 cards per generation

Frequently asked questions

Do I need AwesomeTTS or HyperTTS?

Not for a deck generated here: the mp3 files are already inside the package. An add-on is for adding audio to a deck that does not have any, which is a different problem.

Does the audio work offline?

Yes. The files live in your collection, so reviews work on a plane or on the underground exactly like any other card.

Does audio cost a credit?

No. Audio never consumes a credit, on any plan. One credit is charged for the generation itself, whatever the card count and whether or not you take the audio.

Which languages have audio?

All 20 languages the generator supports, with a neural voice for each. The voice speaks the side in the language you are learning.

Can I add audio to a deck I already have?

Not through NextLang: the audio actions only voice cards from your own results and your own vocabulary, by design. For an existing collection, an add-on inside Anki is the right tool.

Keep reading

Related guides that build on what you just read.