Related terms
A metric that measures whether an AI system's output is factually accurate, logically valid, and satisfies the intended task or expected result.
RelevanceA performance metric that measures how pertinent and applicable generated outputs or retrieved results are to the query.
HelpfulnessAn evaluation metric that measures how effectively an AI system's response addresses the user's request by providing relevant, accurate, complete, and actionable information that satisfies the intended task or objective.
GroundednessAn evaluation metric that measures whether an AI system's output is supported by the provided context, retrieved information, or source material without introducing unsupported claims or hallucinations.
Reference-based EvaluationAn evaluation method that measures model output quality by comparing generated responses against ground-truth reference data or golden datasets.