← Back to modules

Tokenization

BPE optimality and tie-breaking, WordPiece PMI, Unigram EM and Viterbi, subword regularization, byte-to-unicode mapping, glitch tokens, Unicode/homoglyph attacks, character coverage, and tokenizer-free models.

Advanced30 questions
Free account

Take the full module

These are the first few of 30 questions. A free account opens the rest as a scored drill.

  • Every question in this module
  • Instant feedback and supporting reading
  • Your score and progress, saved

Free · your email is used for progress only.