Scenario-based Evaluation
An evaluation method that tests AI performance and behavior across realistic, multi-step scenarios or simulated environments.
Explore more about Evaluation Methods
Related terms
An evaluation method in which human reviewers assess the quality of AI system outputs against defined criteria such as correctness, relevance, helpfulness, safety, or fluency, providing judgments that complement or validate automated evaluation metrics.
Automated EvaluationAn evaluation method that uses software, benchmarks, metrics, or models to assess the quality, correctness, or performance of AI systems without manual review.
Task-specific EvaluationAn assessment method that measures system performance on targeted, custom domain tasks rather than general capabilities.
Offline EvaluationAn evaluation performed outside production using fixed datasets, test cases, or recorded interactions to assess an AI system before deployment.
Rubric-based EvaluationAn evaluation approach that assesses AI outputs against structured, criteria-specific scoring rubrics.