All courses

Guided Course

understandtheorycode

Attention Mechanisms, from First Principles

Self-attention, multi-head, KV cache, MQA, GQA, MLA, FlashAttention — the complete mechanism that makes transformers work.

0/16
Start course
16 lessons
  1. 1. Why attention exists8 min
  2. 2. Scaled dot-product attention8 min
  3. 3. Multi-head attention8 min
  4. 4. Self-attention, cross-attention, and causal masking10 min
  5. 5. Positional encoding — injecting order into attention8 min
  6. 6. RoPE, ALiBi, and long-context position8 min
  7. 7. Autoregressive inference and the KV cache8 min
  8. 8. KV cache memory math8 min
  9. 9. KV cache management — prefix caching, paging, and eviction8 min
  10. 10. Multi-Query Attention (MQA)8 min
  11. 11. Grouped Query Attention (GQA)8 min
  12. 12. Multi-Latent Attention (MLA)8 min
  13. 13. FlashAttention10 min
  14. 14. Sparse and sliding-window attention8 min
  15. 15. Linear attention and state-space models10 min
  16. 16. Choosing the right attention configuration11 min