Comparative Evaluation
An evaluation method that compares the performance of two or more models, systems, or approaches using the same tasks and criteria.
Explore more about Evaluation Methods
Related terms
An evaluation method that compares two outputs directly to determine which performs better against a defined criterion or preference.
Pointwise EvaluationAn evaluation method that assigns an independent score to each individual model response using defined criteria, without comparing it directly to another response.
Blind EvaluationAn evaluation method where evaluators assess outputs without knowing which model, system, or approach produced them to reduce bias.
Model ComparisonAn evaluation method that compares AI models on the same tasks, datasets, or metrics to identify differences in capabilities and performance.
Human EvaluationAn evaluation method in which human reviewers assess the quality of AI system outputs against defined criteria such as correctness, relevance, helpfulness, safety, or fluency, providing judgments that complement or validate automated evaluation metrics.