Nybble™
SnackStackHack
The Transformer Architecture

Lesson 15 of 16

Lessons

  1. 1. Before transformers — the sequential bottleneck
  2. 2. Attention Is All You Need — the transformer
  3. 3. The embedding layer — from token IDs to vectors
  4. 4. The feed-forward network — per-token computation
  5. 5. Normalization — keeping activations in a trainable range
  6. 6. Residual connections and the residual stream
  7. 7. The transformer block and the full stack
  8. 8. Encoder-decoder vs. decoder-only
  9. 9. Mixture of Experts (MoE) — conditional computation at scale
  10. 10. Modern architecture choices — the 2026 consensus
  11. 11. Scaling laws — predicting loss from compute, data, and parameters
  12. 12. Pre-training — learning everything from next-token prediction
  13. 13. Post-training — SFT, RLHF, DPO
  14. 14. Inference — prefill, decode, and sampling
  15. 15. The model landscape and choosing
  16. 16. Architecture decisions for production

The model landscape and choosing

Free account

Read this lesson

The opening lessons of every course are free — this one needs an account. Sign in and the full course opens.

  • Every lesson, start to finish
  • Progress saved across devices
  • Bits per lesson, plus a bonus for finishing

Free · your email is used for progress only.

Nybble™ — built for the people building AI.

PricingAboutRefundsTermsPrivacyContact
SnackStackHack
Message Nybble