All courses

Guided Course

shipcodearchitecture

LLM Evaluation and Testing

llm-evaluation-and-testing

0/27
Start course
27 lessons
  1. 1. Why evaluation is different for AI7 min
  2. 2. Metrics that matter7 min
  3. 3. Choosing metrics for what you actually built7 min
  4. 4. Single-turn scoring: the unit of evaluation6 min
  5. 5. Consistency: the score that moves when nothing changed8 min
  6. 6. Golden datasets8 min
  7. 7. Building a golden set that survives contact6 min
  8. 8. Human annotation and agreement6 min
  9. 9. LLM-as-a-judge8 min
  10. 10. Judge designs: pointwise, pairwise, reference-guided6 min
  11. 11. Judge bias, measured and mitigated5 min
  12. 12. Calibrating and meta-evaluating a judge6 min
  13. 13. Multi-turn evaluation7 min
  14. 14. User simulators7 min
  15. 15. The RAG evaluation triad10 min
  16. 16. Component-level evaluation7 min
  17. 17. Agent evaluation7 min
  18. 18. Regression testing for AI8 min
  19. 19. Slices, peeking, and the multiple-comparisons trap6 min
  20. 20. Building an eval harness7 min
  21. 21. Reproducible runs: the eval manifest6 min
  22. 22. CI/CD gates for AI quality8 min
  23. 23. When offline and online disagree5 min
  24. 24. A/B testing in production7 min
  25. 25. Online monitoring and drift detection8 min
  26. 26. Bias, safety, and adversarial evaluation9 min
  27. 27. Capstone — an evaluation pipeline10 min
  28. Key takeaways4 min
  29. How to get certified3 min
  30. Your certificate