All courses
Guided Course
understandtheoryarchitecture
The Transformer Architecture
From 'Attention Is All You Need' to GPT-4 and Llama 3 — every layer, every design choice, every scaling decision.
0/16
16 lessons
- 1. Before transformers — the sequential bottleneck8 min
- 2. Attention Is All You Need — the transformer8 min
- 3. The embedding layer — from token IDs to vectors8 min
- 4. The feed-forward network — per-token computation8 min
- 5. Normalization — keeping activations in a trainable range9 min
- 6. Residual connections and the residual stream9 min
- 7. The transformer block and the full stack9 min
- 8. Encoder-decoder vs. decoder-only10 min
- 9. Mixture of Experts (MoE) — conditional computation at scale10 min
- 10. Modern architecture choices — the 2026 consensus7 min
- 11. Scaling laws — predicting loss from compute, data, and parameters7 min
- 12. Pre-training — learning everything from next-token prediction10 min
- 13. Post-training — SFT, RLHF, DPO7 min
- 14. Inference — prefill, decode, and sampling8 min
- 15. The model landscape and choosing8 min
- 16. Architecture decisions for production8 min
About Nybble™
The AI space moves fast.
Nybble™ is how you keep up — and stay sharp.
Snack
What happened. In two minutes.
The AI news cycle moves at a pace no one can keep up with. Snack distills what launched, what shipped, and what matters — every day, without the filler.
Go to SnackStack
The concepts behind the headlines.
News tells you what. Stack tells you why and how. From RAG architectures to agentic evals, these are the ideas that will shape what you build next.
Go to StackHack
Prove you actually get it.
Reading about LangChain is not the same as knowing it. Hack challenges you with production-grade questions, then shows you the references that make the answer stick.
Go to Hack