Nybble™
SnackStackHack
LLM Evaluation and Testing

Course wrap-up

Lessons

  1. 1. Why evaluation is different for AI
  2. 2. Metrics that matter
  3. 3. Golden datasets
  4. 4. LLM-as-a-judge
  5. 5. The RAG evaluation triad
  6. 6. Component-level evaluation
  7. 7. Agent evaluation
  8. 8. Regression testing for AI
  9. 9. Building an eval harness
  10. 10. CI/CD gates for AI quality
  11. 11. A/B testing in production
  12. 12. Online monitoring and drift detection
  13. 13. Bias, safety, and adversarial evaluation
  14. 14. Capstone — an evaluation pipeline
  15. Key takeaways
  16. How to get certified
  17. Your certificate

Your certificate

Certificates are issued to an account — it's what ties the credential to a name a reader can check. Sign in to see where you stand on LLM Evaluation and Testing.

← PreviousBack to course

Nybble™ — built for the people building AI.

AboutTermsPrivacyContact
SnackStackHack
Message Nybble