- Published:
- Last updated:
Evaluating AI Applications
- Authors
- Name
- Lucian Oprea
- @LucianDSA_
00:30:00
Filter by difficulty
31 of 31 questions shown
Evaluation Foundations and Success Criteria
⏷ 1. What is an AI application evaluation, and why are normal unit tests not enough?
⏷ 2. How do you define an evaluation objective before choosing metrics?
⏷ 3. What should one evaluation case contain?
⏷ 4. What is the difference between offline and online evaluation?
⏷ 5. Why should evaluations be task-specific instead of relying on one generic quality score?
Free preview complete
You’ve reached the end of the free preview
Get every remaining question and complete answer, plus progress tracking across the full Interview Question Library.
- 26 more questions and complete answers in this topic
- Full access to every interview topic
- Progress tracking and question flags
- New questions and improvements during your subscription
Full access from
$12/month
No long-term commitment. Cancel whenever you want.
⏷ 6. How do you build a useful failure taxonomy for an AI feature?
Datasets, Labels, and Reproducibility
⏷ 7. What makes an evaluation dataset representative?
⏷ 8. What is a golden dataset, and how should it be used?
⏷ 9. How do you split evaluation data and reduce contamination?
⏷ 10. When is synthetic evaluation data useful, and what can go wrong?
⏷ 11. How do you write a labeling rubric that humans can apply consistently?
⏷ 12. How do you version datasets and make evaluation runs reproducible?
Graders, Comparison, and Calibration
⏷ 13. Which checks should be deterministic before using a model grader?
⏷ 14. When should you use an LLM as a judge?
⏷ 15. When is pairwise comparison better than scoring one answer?
⏷ 16. How do you calibrate an automated grader against human judgment?
⏷ 17. How do bias and variance affect AI evaluation results?
⏷ 18. What is grader hacking, and how do you detect it?
RAG and End-to-End System Evaluation
⏷ 19. Why should RAG retrieval and generation be evaluated separately?
⏷ 20. Which metrics help evaluate ranked retrieval results?
⏷ 21. How do groundedness, relevance, and completeness differ in a RAG evaluation?
⏷ 22. How do you evaluate chunking, top-k, and reranking changes?
⏷ 23. How do you evaluate agents with tools without grading only the final answer?
⏷ 24. How do you include latency, cost, and reliability in evaluation gates?
Regression, Experiments, and Production Feedback
⏷ 25. What belongs in a CI regression suite for an AI feature?
⏷ 26. How do you compare prompt, model, or configuration changes fairly?
⏷ 27. How do you set release thresholds without hiding slice regressions?
⏷ 28. How do you turn production feedback into better evaluations?
⏷ 29. How do you detect data, model, and behavior drift in production?
⏷ 30. How do you design a safe canary or A/B test for an AI change?
Statistical Confidence
⏷ 31. How do sample size, confidence intervals, and repeated trials affect an AI evaluation?