TruthfulQA
A benchmark measuring how truthfully language models answer questions and avoid mimicking human falsehoods.
Explore more about Benchmarks
Related terms
A benchmark for evaluating language models across diverse subjects spanning humanities, social sciences, STEM, and professional knowledge.
MMLU-ProA more challenging version of MMLU designed to evaluate advanced reasoning and knowledge across diverse academic and professional subjects.
GPQAGPQAA benchmark that evaluates the ability of AI models to answer challenging graduate-level multiple-choice questions across scientific domains, measuring expert-level reasoning, scientific knowledge, and problem-solving performance.
ARCARCA benchmark that evaluates an AI model's ability to solve grade-school science questions requiring reasoning, knowledge, and commonsense understanding.
HellaSwagA benchmark that evaluates the commonsense reasoning and natural language understanding capabilities of AI models by requiring them to select the most plausible continuation of a given context from multiple candidate endings.