Related terms
A metric that evaluates generated text by measuring n-gram overlap between a candidate output and one or more reference texts.
ROUGEROUGEA set of evaluation metrics used to measure the overlap of n-grams between generated text and reference summaries.
METEORMETEORA metric for evaluating generated text by measuring alignment with reference text using word matches, stemming, and synonym matching.
Reference-based EvaluationAn evaluation method that measures model output quality by comparing generated responses against ground-truth reference data or golden datasets.
RelevanceA performance metric that measures how pertinent and applicable generated outputs or retrieved results are to the query.