Evaluation Dataset
Also called: Evaluation Set, Test Dataset
A curated collection of test examples, inputs, and expected outcomes used to measure the quality, accuracy, safety, and reliability of AI models and systems.
Explore more about Feedback Loops
Related terms
A curated and validated collection of high-quality reference examples with trusted labels or expected outputs that serves as a benchmark for evaluating, testing, and monitoring the performance of AI models and applications.
Reference-based EvaluationAn evaluation method that measures model output quality by comparing generated responses against ground-truth reference data or golden datasets.
Human EvaluationAn evaluation method in which human reviewers assess the quality of AI system outputs against defined criteria such as correctness, relevance, helpfulness, safety, or fluency, providing judgments that complement or validate automated evaluation metrics.
Data CurationThe process of collecting, organizing, cleaning, validating, and maintaining datasets to improve the quality of AI training, evaluation, and feedback pipelines.