All courses
Guided Course
understandtheoryarchitecture
The Transformer Architecture
From 'Attention Is All You Need' to GPT-4 and Llama 3 — every layer, every design choice, every scaling decision.
0/16
16 lessons
- 1. Before transformers — the sequential bottleneck11 min
- 2. Attention Is All You Need — the transformer10 min
- 3. The embedding layer — from token IDs to vectors11 min
- 4. The feed-forward network — per-token computation10 min
- 5. Normalization — keeping activations in a trainable range10 min
- 6. Residual connections and the residual stream11 min
- 7. The transformer block and the full stack12 min
- 8. Encoder-decoder vs. decoder-only10 min
- 9. Mixture of Experts (MoE) — conditional computation at scale10 min
- 10. Modern architecture choices — the 2026 consensus8 min
- 11. Scaling laws — predicting loss from compute, data, and parameters7 min
- 12. Pre-training — learning everything from next-token prediction15 min
- 13. Post-training — SFT, RLHF, DPO7 min
- 14. Inference — prefill, decode, and sampling10 min
- 15. The model landscape and choosing8 min
- 16. Architecture decisions for production8 min
- Key takeaways4 min
- How to get certified3 min
- Your certificate
About Nybble™
The AI space moves fast.
Nybble™ is how you keep up — and stay sharp.
Snack
What happened. In two minutes.
The AI news cycle moves at a pace no one can keep up with. Snack distills what launched, what shipped, and what matters — every day, without the filler.
Go to SnackStack
The concepts behind the headlines.
News tells you what. Stack tells you why and how. From RAG architectures to agentic evals, these are the ideas that will shape what you build next.
Go to StackHack
Prove you actually get it.
Reading about LangChain is not the same as knowing it. Hack challenges you with production-grade questions, then shows you the references that make the answer stick.
Go to Hack