Nybble™
SnackStackHack
Tokenization, Inside and Out

Course wrap-up

Lessons

  1. 1. From text to numbers
  2. 2. The tokenizer pipeline
  3. 3. Byte Pair Encoding (BPE)
  4. 4. Byte-level BPE
  5. 5. WordPiece — likelihood-driven subword merging
  6. 6. Unigram — top-down probabilistic segmentation
  7. 7. SentencePiece and tiktoken — the two dominant tokenizer implementations
  8. 8. Vocabulary design
  9. 9. Special tokens and chat templates
  10. 10. Multilingual tokenization and fairness
  11. 11. Token counting and cost estimation
  12. 12. Tokenizer fragility and adversarial inputs
  13. 13. Training a tokenizer from scratch
  14. 14. The tokenizer-free future
  15. Key takeaways
  16. How to get certified
  17. Your certificate

Your certificate

Certificates are issued to an account — it's what ties the credential to a name a reader can check. Sign in to see where you stand on Tokenization, Inside and Out.

← PreviousBack to course

Nybble™ — built for the people building AI.

AboutTermsPrivacyContact
SnackStackHack
Message Nybble