Nybble™
SnackStackHack
Attention Mechanisms, from First Principles

Lesson 13 of 16

Lessons

  1. 1. Why attention exists
  2. 2. Scaled dot-product attention
  3. 3. Multi-head attention
  4. 4. Self-attention, cross-attention, and causal masking
  5. 5. Positional encoding — injecting order into attention
  6. 6. RoPE, ALiBi, and long-context position
  7. 7. Autoregressive inference and the KV cache
  8. 8. KV cache memory math
  9. 9. KV cache management — prefix caching, paging, and eviction
  10. 10. Multi-Query Attention (MQA)
  11. 11. Grouped Query Attention (GQA)
  12. 12. Multi-Latent Attention (MLA)
  13. 13. FlashAttention
  14. 14. Sparse and sliding-window attention
  15. 15. Linear attention and state-space models
  16. 16. Choosing the right attention configuration

FlashAttention

Free account

Read this lesson

The opening lessons of every course are free — this one needs an account. Sign in and the full course opens.

  • Every lesson, start to finish
  • Progress saved across devices
  • Bits per lesson, plus a bonus for finishing

Free · your email is used for progress only.

Nybble™ — built for the people building AI.

PricingAboutRefundsTermsPrivacyContact
SnackStackHack
Message Nybble