LLM-as-a-Judge
An evaluation method that uses a large language model to assess the quality, correctness, or other attributes of AI-generated outputs.
Explore more about Evaluation Methods
Related terms
An evaluation approach that uses another trained model to assess the quality, correctness, safety, or other properties of an AI system's outputs.
Pairwise ComparisonAn evaluation method that compares two outputs directly to determine which performs better against a defined criterion or preference.
Pointwise EvaluationAn evaluation method that assigns an independent score to each individual model response using defined criteria, without comparing it directly to another response.
Rubric-based EvaluationAn evaluation approach that assesses AI outputs against structured, criteria-specific scoring rubrics.