All courses

Guided Course

understandtheoryarchitecture

The Transformer Architecture

From 'Attention Is All You Need' to GPT-4 and Llama 3 — every layer, every design choice, every scaling decision.

0/16
Start course
16 lessons
  1. 1. Before transformers — the sequential bottleneck8 min
  2. 2. Attention Is All You Need — the transformer8 min
  3. 3. The embedding layer — from token IDs to vectors8 min
  4. 4. The feed-forward network — per-token computation8 min
  5. 5. Normalization — keeping activations in a trainable range9 min
  6. 6. Residual connections and the residual stream9 min
  7. 7. The transformer block and the full stack9 min
  8. 8. Encoder-decoder vs. decoder-only10 min
  9. 9. Mixture of Experts (MoE) — conditional computation at scale10 min
  10. 10. Modern architecture choices — the 2026 consensus7 min
  11. 11. Scaling laws — predicting loss from compute, data, and parameters7 min
  12. 12. Pre-training — learning everything from next-token prediction10 min
  13. 13. Post-training — SFT, RLHF, DPO7 min
  14. 14. Inference — prefill, decode, and sampling8 min
  15. 15. The model landscape and choosing8 min
  16. 16. Architecture decisions for production8 min