All courses
Guided Course
shipcodearchitecture
LLM Evaluation and Testing
llm-evaluation-and-testing
0/27
27 lessons
- 1. Why evaluation is different for AI7 min
- 2. Metrics that matter7 min
- 3. Choosing metrics for what you actually built7 min
- 4. Single-turn scoring: the unit of evaluation6 min
- 5. Consistency: the score that moves when nothing changed8 min
- 6. Golden datasets8 min
- 7. Building a golden set that survives contact6 min
- 8. Human annotation and agreement6 min
- 9. LLM-as-a-judge8 min
- 10. Judge designs: pointwise, pairwise, reference-guided6 min
- 11. Judge bias, measured and mitigated5 min
- 12. Calibrating and meta-evaluating a judge6 min
- 13. Multi-turn evaluation7 min
- 14. User simulators7 min
- 15. The RAG evaluation triad10 min
- 16. Component-level evaluation7 min
- 17. Agent evaluation7 min
- 18. Regression testing for AI8 min
- 19. Slices, peeking, and the multiple-comparisons trap6 min
- 20. Building an eval harness7 min
- 21. Reproducible runs: the eval manifest6 min
- 22. CI/CD gates for AI quality8 min
- 23. When offline and online disagree5 min
- 24. A/B testing in production7 min
- 25. Online monitoring and drift detection8 min
- 26. Bias, safety, and adversarial evaluation9 min
- 27. Capstone — an evaluation pipeline10 min
- Key takeaways4 min
- How to get certified3 min
- Your certificate
About Nybble™
The AI space moves fast.
Nybble™ is how you keep up — and stay sharp.
Snack
What happened. In two minutes.
The AI news cycle moves at a pace no one can keep up with. Snack distills what launched, what shipped, and what matters — every day, without the filler.
Go to SnackStack
The concepts behind the headlines.
News tells you what. Stack tells you why and how. From RAG architectures to agentic evals, these are the ideas that will shape what you build next.
Go to StackHack
Prove you actually get it.
Reading about LangChain is not the same as knowing it. Hack challenges you with production-grade questions, then shows you the references that make the answer stick.
Go to Hack