Related terms
An evaluation method that uses a large language model to assess the quality, correctness, or other attributes of AI-generated outputs.
Continuous EvaluationAn evaluation approach that continuously measures AI system performance throughout development and production to detect regressions and ensure quality.
Human EvaluationAn evaluation method in which human reviewers assess the quality of AI system outputs against defined criteria such as correctness, relevance, helpfulness, safety, or fluency, providing judgments that complement or validate automated evaluation metrics.
Reference-based EvaluationAn evaluation method that measures model output quality by comparing generated responses against ground-truth reference data or golden datasets.
AI AgentA software system that perceives its environment, reasons about goals, and autonomously performs actions using AI models, tools, memory, and workflows.