Nybble™
SnackStackHack
The Transformer Architecture

Lesson 11 of 16

Lessons

  1. 1. Before transformers — the sequential bottleneck
  2. 2. Attention Is All You Need — the transformer
  3. 3. The embedding layer — from token IDs to vectors
  4. 4. The feed-forward network — per-token computation
  5. 5. Normalization — keeping activations in a trainable range
  6. 6. Residual connections and the residual stream
  7. 7. The transformer block and the full stack
  8. 8. Encoder-decoder vs. decoder-only
  9. 9. Mixture of Experts (MoE) — conditional computation at scale
  10. 10. Modern architecture choices — the 2026 consensus
  11. 11. Scaling laws — predicting loss from compute, data, and parameters
  12. 12. Pre-training — learning everything from next-token prediction
  13. 13. Post-training — SFT, RLHF, DPO
  14. 14. Inference — prefill, decode, and sampling
  15. 15. The model landscape and choosing
  16. 16. Architecture decisions for production

Scaling laws — predicting loss from compute, data, and parameters

Free account

Read this lesson

The opening lessons of every course are free — this one needs an account. Sign in and the full course opens.

  • Every lesson, start to finish
  • Progress saved across devices
  • Bits per lesson, plus a bonus for finishing

Free · your email is used for progress only.

Nybble™ — built for the people building AI.

PricingAboutRefundsTermsPrivacyContact
SnackStackHack
Message Nybble