MMLU-Pro
A more challenging version of MMLU designed to evaluate advanced reasoning and knowledge across diverse academic and professional subjects.
Explore more about Benchmarks
Related terms
A benchmark for evaluating language models across diverse subjects spanning humanities, social sciences, STEM, and professional knowledge.
GPQAGPQAA benchmark that evaluates the ability of AI models to answer challenging graduate-level multiple-choice questions across scientific domains, measuring expert-level reasoning, scientific knowledge, and problem-solving performance.
MATHMATHA benchmark for evaluating mathematical problem-solving abilities of AI models across challenging competition-level mathematics problems.
Benchmark RunA single execution of a benchmark used to measure and record a model or system's performance under defined conditions.