Task-specific Evaluation
An assessment method that measures system performance on targeted, custom domain tasks rather than general capabilities.
Explore more about Evaluation Methods
Related terms
An evaluation method in which human reviewers assess the quality of AI system outputs against defined criteria such as correctness, relevance, helpfulness, safety, or fluency, providing judgments that complement or validate automated evaluation metrics.
Automated EvaluationAn evaluation method that uses software, benchmarks, metrics, or models to assess the quality, correctness, or performance of AI systems without manual review.
Model-based EvaluationAn evaluation approach that uses another trained model to assess the quality, correctness, safety, or other properties of an AI system's outputs.
Rubric-based EvaluationAn evaluation approach that assesses AI outputs against structured, criteria-specific scoring rubrics.
Scenario-based EvaluationAn evaluation method that tests AI performance and behavior across realistic, multi-step scenarios or simulated environments.