← Back to modules

LLM Training Pipeline

Chinchilla vs Kaplan, compute vs inference optimality, AdamW internals, 3D parallelism, ZeRO stages, stability/loss spikes, MFU, reward modeling, PPO vs DPO, RLAIF, reward hacking.

Advanced50 questions
Free account

Take the full module

These are the first few of 50 questions. A free account opens the rest as a scored drill.

  • Every question in this module
  • Instant feedback and supporting reading
  • Your score and progress, saved

Free · your email is used for progress only.