User Study
An evaluation method that assesses system performance and usability through direct observation and structured human interaction.
Explore more about Evaluation Methods
Related terms
An evaluation method in which human reviewers assess the quality of AI system outputs against defined criteria such as correctness, relevance, helpfulness, safety, or fluency, providing judgments that complement or validate automated evaluation metrics.
Expert ReviewAn evaluation method in which subject matter experts assess the quality, accuracy, safety, or effectiveness of an AI system, model, or output using their domain knowledge and established criteria.
Blind EvaluationAn evaluation method where evaluators assess outputs without knowing which model, system, or approach produced them to reduce bias.
Acceptance TestingA testing process that verifies whether an AI system satisfies specified requirements and is ready for deployment or release.
Scenario-based EvaluationAn evaluation method that tests AI performance and behavior across realistic, multi-step scenarios or simulated environments.