Related terms
A metric that measures the proportion of correct predictions or outputs produced by a model or AI system out of all evaluated cases.
Pass@kA code-generation metric measuring the probability that at least one of k generated solutions correctly solves a given programming problem.
Win RateThe proportion of evaluation tasks or head-to-head comparisons in which an AI model outperforms a baseline or competitor.
CorrectnessA metric that measures whether an AI system's output is factually accurate, logically valid, and satisfies the intended task or expected result.
ReliabilityA performance metric measuring the degree to which a system consistently performs its intended function without failure over time.