← Back to modules

Transformer Architecture

Attention complexity and FlashAttention, MHA vs MQA vs GQA, sparse and sliding-window attention, KV-cache math, Pre/Post-Norm and DeepNorm, attention sinks, head redundancy, and training-stability trade-offs.

Advanced50 questions
Free account

Take the full module

These are the first few of 50 questions. A free account opens the rest as a scored drill.

  • Every question in this module
  • Instant feedback and supporting reading
  • Your score and progress, saved

Free · your email is used for progress only.