Nybble™
SnackStackHack
Tokenization, Inside and Out

Lesson 3 of 14

Lessons

  1. 1. From text to numbers
  2. 2. The tokenizer pipeline
  3. 3. Byte Pair Encoding (BPE)
  4. 4. Byte-level BPE
  5. 5. WordPiece — likelihood-driven subword merging
  6. 6. Unigram — top-down probabilistic segmentation
  7. 7. SentencePiece and tiktoken — the two dominant tokenizer implementations
  8. 8. Vocabulary design
  9. 9. Special tokens and chat templates
  10. 10. Multilingual tokenization and fairness
  11. 11. Token counting and cost estimation
  12. 12. Tokenizer fragility and adversarial inputs
  13. 13. Training a tokenizer from scratch
  14. 14. The tokenizer-free future

Byte Pair Encoding (BPE)

Free account

Read this lesson

The opening lessons of every course are free — this one needs an account. Sign in and the full course opens.

  • Every lesson, start to finish
  • Progress saved across devices
  • Bits per lesson, plus a bonus for finishing

Free · your email is used for progress only.

Nybble™ — built for the people building AI.

PricingAboutRefundsTermsPrivacyContact
SnackStackHack
Message Nybble