Related terms
A metric that measures the proportion of correct predictions or outputs produced by a model or AI system out of all evaluated cases.
Success RateThe proportion of attempted tasks or requests that an AI agent or system completes successfully.
Pass@kA code-generation metric measuring the probability that at least one of k generated solutions correctly solves a given programming problem.
Pairwise ComparisonAn evaluation method that compares two outputs directly to determine which performs better against a defined criterion or preference.
LLM-as-a-JudgeAn evaluation method that uses a large language model to assess the quality, correctness, or other attributes of AI-generated outputs.