All courses

Guided Course

shipopsarchitecture

LLMOps: Serving and Operating LLMs

Model serving, inference optimization, guardrails, observability, and cost control — the engineering that keeps LLM systems running.

0/17
Start course
17 lessons
  1. 1. The model serving stack7 min
  2. 2. vLLM and production model servers9 min
  3. 3. Inference optimization10 min
  4. 4. Model compression: PTQ, QAT, and calibration7 min
  5. 5. Pruning and distillation6 min
  6. 6. Streaming and real-time delivery8 min
  7. 7. Building LLM APIs8 min
  8. 8. Model routing and fallbacks9 min
  9. 9. Caching strategies9 min
  10. 10. Input guardrails8 min
  11. 11. Output guardrails9 min
  12. 12. The OWASP Top 10 for LLM applications10 min
  13. 13. LLM observability7 min
  14. 14. Cost control7 min
  15. 15. Capacity planning and scaling7 min
  16. 16. Incident response and runbooks7 min
  17. 17. Capstone — a production LLM service9 min
  18. Key takeaways4 min
  19. How to get certified3 min
  20. Your certificate