← Back to modules

Quantization & Compression

Method internals and trade-offs: GPTQ/AWQ/SmoothQuant/LLM.int8, NF4 and double quant, FP8, SparseGPT vs Wanda, KV-cache quant, extreme low-bit, distillation objectives.

Advanced50 questions
Free account

Take the full module

These are the first few of 50 questions. A free account opens the rest as a scored drill.

  • Every question in this module
  • Instant feedback and supporting reading
  • Your score and progress, saved

Free · your email is used for progress only.