Nybble™
SnackStackHack
LLMOps: Serving and Operating LLMs

Lesson 17 of 17

Lessons

  1. 1. The model serving stack
  2. 2. vLLM and production model servers
  3. 3. Inference optimization
  4. 4. Model compression: PTQ, QAT, and calibration
  5. 5. Pruning and distillation
  6. 6. Streaming and real-time delivery
  7. 7. Building LLM APIs
  8. 8. Model routing and fallbacks
  9. 9. Caching strategies
  10. 10. Input guardrails
  11. 11. Output guardrails
  12. 12. The OWASP Top 10 for LLM applications
  13. 13. LLM observability
  14. 14. Cost control
  15. 15. Capacity planning and scaling
  16. 16. Incident response and runbooks
  17. 17. Capstone — a production LLM service
  18. Key takeaways
  19. How to get certified
  20. Your certificate

Capstone — a production LLM service

Free account

Read this lesson

The opening lessons of every course are free — this one needs an account. Sign in and the full course opens.

  • Every lesson, start to finish
  • Progress saved across devices
  • Bits per lesson, plus a bonus for finishing

Free · your email is used for progress only.

Nybble™ — built for the people building AI.

AboutTermsPrivacyContact
SnackStackHack
Message Nybble