Blind Evaluation
Also called: Blind Testing
An evaluation method where evaluators assess outputs without knowing which model, system, or approach produced them to reduce bias.
Explore more about Evaluation Methods
Related terms
An evaluation method in which human reviewers assess the quality of AI system outputs against defined criteria such as correctness, relevance, helpfulness, safety, or fluency, providing judgments that complement or validate automated evaluation metrics.
Pairwise ComparisonAn evaluation method that compares two outputs directly to determine which performs better against a defined criterion or preference.
Comparative EvaluationAn evaluation method that compares the performance of two or more models, systems, or approaches using the same tasks and criteria.
Expert ReviewAn evaluation method in which subject matter experts assess the quality, accuracy, safety, or effectiveness of an AI system, model, or output using their domain knowledge and established criteria.
User StudyAn evaluation method that assesses system performance and usability through direct observation and structured human interaction.