Nybble™
SnackStackHack
The Transformer Architecture

Course wrap-up

Lessons

  1. 1. Before transformers — the sequential bottleneck
  2. 2. Attention Is All You Need — the transformer
  3. 3. The embedding layer — from token IDs to vectors
  4. 4. The feed-forward network — per-token computation
  5. 5. Normalization — keeping activations in a trainable range
  6. 6. Residual connections and the residual stream
  7. 7. The transformer block and the full stack
  8. 8. Encoder-decoder vs. decoder-only
  9. 9. Mixture of Experts (MoE) — conditional computation at scale
  10. 10. Modern architecture choices — the 2026 consensus
  11. 11. Scaling laws — predicting loss from compute, data, and parameters
  12. 12. Pre-training — learning everything from next-token prediction
  13. 13. Post-training — SFT, RLHF, DPO
  14. 14. Inference — prefill, decode, and sampling
  15. 15. The model landscape and choosing
  16. 16. Architecture decisions for production
  17. Key takeaways
  18. How to get certified
  19. Your certificate

Your certificate

Certificates are issued to an account — it's what ties the credential to a name a reader can check. Sign in to see where you stand on The Transformer Architecture.

← PreviousBack to course

Nybble™ — built for the people building AI.

AboutTermsPrivacyContact
SnackStackHack
Message Nybble