Offline Evaluation
An evaluation performed outside production using fixed datasets, test cases, or recorded interactions to assess an AI system before deployment.
Explore more about Evaluation Methods
Related terms
A curated collection of test examples, inputs, and expected outcomes used to measure the quality, accuracy, safety, and reliability of AI models and systems.
Benchmark RunA single execution of a benchmark used to measure and record a model or system's performance under defined conditions.
Reference-based EvaluationAn evaluation method that measures model output quality by comparing generated responses against ground-truth reference data or golden datasets.
Automated EvaluationAn evaluation method that uses software, benchmarks, metrics, or models to assess the quality, correctness, or performance of AI systems without manual review.