All courses
Guided Course
understandtheorycode
Attention Mechanisms, from First Principles
Self-attention, multi-head, KV cache, MQA, GQA, MLA, FlashAttention — the complete mechanism that makes transformers work.
0/16
16 lessons
- 1. Why attention exists8 min
- 2. Scaled dot-product attention8 min
- 3. Multi-head attention8 min
- 4. Self-attention, cross-attention, and causal masking10 min
- 5. Positional encoding — injecting order into attention8 min
- 6. RoPE, ALiBi, and long-context position8 min
- 7. Autoregressive inference and the KV cache8 min
- 8. KV cache memory math8 min
- 9. KV cache management — prefix caching, paging, and eviction8 min
- 10. Multi-Query Attention (MQA)8 min
- 11. Grouped Query Attention (GQA)8 min
- 12. Multi-Latent Attention (MLA)8 min
- 13. FlashAttention10 min
- 14. Sparse and sliding-window attention8 min
- 15. Linear attention and state-space models10 min
- 16. Choosing the right attention configuration11 min
About Nybble™
The AI space moves fast.
Nybble™ is how you keep up — and stay sharp.
Snack
What happened. In two minutes.
The AI news cycle moves at a pace no one can keep up with. Snack distills what launched, what shipped, and what matters — every day, without the filler.
Go to SnackStack
The concepts behind the headlines.
News tells you what. Stack tells you why and how. From RAG architectures to agentic evals, these are the ideas that will shape what you build next.
Go to StackHack
Prove you actually get it.
Reading about LangChain is not the same as knowing it. Hack challenges you with production-grade questions, then shows you the references that make the answer stick.
Go to Hack