Human Evaluation
Also called: Manual Evaluation, Human Assessment
An evaluation method in which human reviewers assess the quality of AI system outputs against defined criteria such as correctness, relevance, helpfulness, safety, or fluency, providing judgments that complement or validate automated evaluation metrics.
Explore more about Evaluation Methods
Related terms
An evaluation method in which subject matter experts assess the quality, accuracy, safety, or effectiveness of an AI system, model, or output using their domain knowledge and established criteria.
LLM-as-a-JudgeAn evaluation method that uses a large language model to assess the quality, correctness, or other attributes of AI-generated outputs.
Evaluation DatasetA curated collection of test examples, inputs, and expected outcomes used to measure the quality, accuracy, safety, and reliability of AI models and systems.
Reference-based EvaluationAn evaluation method that measures model output quality by comparing generated responses against ground-truth reference data or golden datasets.
HelpfulnessAn evaluation metric that measures how effectively an AI system's response addresses the user's request by providing relevant, accurate, complete, and actionable information that satisfies the intended task or objective.