← Back to modules

Inference Optimization

Roofline reasoning, FlashAttention internals, chunked prefill and disaggregation, PagedAttention internals, speculative-decoding acceptance, KV reduction, and tensor-parallel serving.

Advanced50 questions
Free account

Take the full module

These are the first few of 50 questions. A free account opens the rest as a scored drill.

  • Every question in this module
  • Instant feedback and supporting reading
  • Your score and progress, saved

Free · your email is used for progress only.