All courses

Guided Course

understandtheoryarchitecture

The Transformer Architecture

From 'Attention Is All You Need' to GPT-4 and Llama 3 — every layer, every design choice, every scaling decision.

0/16
Start course
16 lessons
  1. 1. Before transformers — the sequential bottleneck11 min
  2. 2. Attention Is All You Need — the transformer10 min
  3. 3. The embedding layer — from token IDs to vectors11 min
  4. 4. The feed-forward network — per-token computation10 min
  5. 5. Normalization — keeping activations in a trainable range10 min
  6. 6. Residual connections and the residual stream11 min
  7. 7. The transformer block and the full stack12 min
  8. 8. Encoder-decoder vs. decoder-only10 min
  9. 9. Mixture of Experts (MoE) — conditional computation at scale10 min
  10. 10. Modern architecture choices — the 2026 consensus8 min
  11. 11. Scaling laws — predicting loss from compute, data, and parameters7 min
  12. 12. Pre-training — learning everything from next-token prediction15 min
  13. 13. Post-training — SFT, RLHF, DPO7 min
  14. 14. Inference — prefill, decode, and sampling10 min
  15. 15. The model landscape and choosing8 min
  16. 16. Architecture decisions for production8 min
  17. Key takeaways4 min
  18. How to get certified3 min
  19. Your certificate