Nybble™
SnackStackHack
Attention Mechanisms, from First Principles

Course wrap-up

Lessons

  1. 1. Why attention exists
  2. 2. Scaled dot-product attention
  3. 3. Multi-head attention
  4. 4. Self-attention, cross-attention, and causal masking
  5. 5. Positional encoding — injecting order into attention
  6. 6. RoPE, ALiBi, and long-context position
  7. 7. Autoregressive inference and the KV cache
  8. 8. KV cache memory math
  9. 9. KV cache management — prefix caching, paging, and eviction
  10. 10. Multi-Query Attention (MQA)
  11. 11. Grouped Query Attention (GQA)
  12. 12. Multi-Latent Attention (MLA)
  13. 13. FlashAttention
  14. 14. Sparse and sliding-window attention
  15. 15. Linear attention and state-space models
  16. 16. Choosing the right attention configuration
  17. Key takeaways
  18. How to get certified
  19. Your certificate

Your certificate

Certificates are issued to an account — it's what ties the credential to a name a reader can check. Sign in to see where you stand on Attention Mechanisms, from First Principles.

← PreviousBack to course

Nybble™ — built for the people building AI.

AboutTermsPrivacyContact
SnackStackHack
Message Nybble