A script to create and manage my own printable kanji flashcards.
Typst 74%
Python 20%
Nix 6%

README.md

Japanese Flashcards #

Build printable JLPT kanji flashcards from raw kanji lists. A single utility (deck.py) drives the whole pipeline: it turns pending.txt (one kanji per line) into full deck entries — readings, meanings, JLPT category and examples all fetched from jisho.org — then generates a Typst document that compiles to a PDF meant for duplex (flip-on-long-edge) printing. Each card shows the bare kanji on the front — the prompt you recall against — and on the back the readings, stroke-order reference, English meaning and example usage.

Sample card from the deck — front and back

Quick start #

Put the two fonts in fonts/ first (see Fonts), then:

nix develop           # recommended: Python 3.10+, typst, font path preconfigured
python3 deck.py generate

deck.py generate renders deck.json to flashcards.typ and compiles it to flashcards.pdf — 10 cards per sheet, laid out so a duplex (flip-on-long-edge) print comes out right. Want the hands-on details instead of a one-liner? See Adding kanji and Regenerating the PDF.

Requirements #

  • Python 3.10+ (standard library only — no third-party packages).
  • Typst ≥ 0.15 to compile the PDF. generate shells out to typst; if it is not on PATH, the .typ is still written.
  • The two fonts in fonts/ — see Fonts.
nix develop   # python312, typst, tinymist, plus a font-specimen helper

The dev-shell exports TYPST_FONT_PATHS=./fonts, so typst compile picks up the deck fonts automatically.

Manual setup #

Install typst your usual way, then always pass the font directory:

typst compile --font-path fonts flashcards.typ

fonts/ must contain a font whose family is Noto Sans JP and one whose family is KanjiStrokeOrders — see Fonts below.

Fonts #

Both fonts are freely licensed but not distributed in the repository, so they are git-ignored and you download them yourself. Copy the two files into fonts/ (their file names don't matter; Typst matches on the embedded family name).

Noto Sans JP — SIL Open Font License 1.1. Used for the main card text. Download the family from Google Fonts ("Get font"), or the 16_NotoSansJP.zip asset from the googlefonts/noto-cjk releases and use static/NotoSansJP-Regular.ttf.

Kanji Stroke Order Font v4.005 — BSD-style license (© Ulrich Apel, the AAAA project and the Wadoku project). Used for the stroke-order rendering on the back of each card. Download KanjiStrokeOrders_v4.005.ttf from kanji.uk (also hosts the full package with the README and licence text).

Adding kanji #

  1. Add the new characters to pending.txt (one per line).
  2. Run python3 deck.py import pending.txt — new entries are built automatically (readings, meaning, JLPT category and examples fetched from jisho.org) and deck.json is updated in place.
  3. Validate with python3 deck.py check deck.json — 0 ERRORs, no warnings, only explainable MISSING items.
  4. Generate with python3 deck.py generate.

Regenerating the PDF #

python3 deck.py import pending.txt   # only if pending.txt gained characters
python3 deck.py generate

Workflow utility (deck.py) #

deck.py import LISTFILE #

Adds every kanji in a plain-text list (one per line, # and blank lines ignored) to deck.json:

python3 deck.py import pending.txt

For each character it:

  • appends a new entry if the character is not yet in the deck, fetching onyomi/kunyomi, meaning and the JLPT category from the kanji page and filling examples from the word API;
  • completes any existing-but-incomplete entry (only ever fills empty fields — it never overwrites hand-edited data);
  • back-fills examples for complete entries that have none.

1–2 kanji are also good to leave offline-made quibbles alone: a character already present with full data is skipped untouched. import checkpoints the deck every 20 characters.

deck.py generate [FILE] #

Renders a deck JSON to flashcards.typ and compiles it to PDF: 10 cards per sheet (2 columns × 5 rows), facing pages with the bare kanji front (kanji-only) and the full answer card back (kanji-card with readings, stroke-order kanji, meaning and examples), and the two columns swapped on the back so a duplex print flips correctly.

python3 deck.py generate                  # default FILE: deck.json
python3 deck.py generate --no-compile     # write .typ without running typst
# --output NAME (default flashcards), --font-path fonts

Data maintenance commands #

Every command takes a deck JSON as its optional final argument (default deck.json):

Command What it does
check Structural validation plus a cross-check of readings against the examples. Exits non-zero if any ERROR is found.
convert Migrates legacy string readings to arrays (in place): splits on spaces / , / 、, drops lone - placeholders, removes exact duplicates.
fetch-readings Rebuilds onyomi/kunyomi from the kanji's page on jisho.org (the authoritative source the readings were originally scraped from). One HTTP request per kanji, ~0.4 s spacing, checkpoint saves every 20.
fetch-examples Fills the examples field for entries that have none, from the jisho.org word API. Keeps the first 3 results per kanji, waits 0.5 s between requests, checkpoint saves every 10 fetches.

convert, fetch-readings and fetch-examples only ever change their own fields.

check findings:

Level Meaning
ERROR Field is still a string or is not a list. Fails the run (exit 1).
WARN Suspicious token: lone -, unexpected characters, katakana in kunyomi, hiragana in onyomi, or an exact duplicate.
MISSING An example word that is exactly the kanji (e.g. • 田 [た]) uses a reading absent from onyomi/kunyomi.

MISSING is advisory: the example may legitimately use a Chinese/mahjong reading (道 [タオ]), an archaic one (年 [とせ]), or a compound-suffix form (所 [じょ]) that jisho does not list in a kanji's main on/kun lists. Only investigate when the reading is common and clearly absent — the classic case was 田 missing た.

Data format (deck.json) #

A JSON array of entries:

{
  "category": "jlptn5",
  "character": "田",
  "onyomi": ["デン"],
  "kunyomi": ["た"],
  "meaning": "rice field, rice paddy",
  "examples": ["• 田 [た] - rice field"]
}
Field Type Meaning
category string jlptn5 … jlptn1 (the deck spans 80 / 96 / 25 / 14 / 3 entries); empty when the kanji has no JLPT level.
character string The kanji (single character).
onyomi array of strings Chinese-derived readings in katakana.
kunyomi array of strings Native readings in hiragana.
meaning string English meaning.
examples array of strings Up to 3 usage examples as • word [reading] - meaning.

An empty reading list is stored as [] and rendered as × on the card.

Reading-token conventions #

Tokens may carry markers that are meaningful and must survive hand edits:

  • leading - — the reading attaches to a preceding element (-び, -あ.げる)
  • trailing - — the reading attaches to a following element (ひと-, うわ-)
  • . — okurigana boundary (わ.ける = the stem ends at わ)

ひ, -び and ひと- are therefore three distinct readings, not duplicates.

Repository layout #

Path Purpose
deck.py End-to-end workflow utility (see Workflow utility).
deck.json The deck. Every kanji with readings, meaning, JLPT category and examples.
pending.txt One kanji per line — the characters still to be reviewed/added.
flashcard-layout.typ Typst layout functions (kanji-card, kanji-only, fmt-readings).
flashcards.typ / flashcards.pdf Generated output (both git-ignored).
preview.png Sample card image shown at the top of this README.
fonts/ Fonts used at compile time — downloaded locally, git-ignored (see Fonts).
flake.nix / flake.lock Nix dev-shell with Python, Typst and the font path configured.

Pipeline overview #

pending.txt                    one kanji per line (the new characters)
   │
   │  deck.py import pending.txt         fetch what's missing, merge into deck.json
   ▼
deck.json (the deck)
   │
   │  deck.py generate [FILE]            default input: deck.json
   ▼
flashcards.typ  ──── typst ────► flashcards.pdf

deck.py edits JSON files in place and every command is idempotent, so a failed run can simply be re-run. After any manual change to readings, run deck.py check deck.json to confirm the file is still consistent.

Gotchas #

  • flashcards.typ is overwritten on every run; flashcards.pdf is git-ignored.
  • Reading markers (-, .) are meaningful — see "Reading-token conventions".
  • An empty reading list renders as ×; don't re-add a - placeholder unless it reflects an actual "attaches to previous/next" form.