MMLU
MMLUAlso called: Massive Multitask Language Understanding
A benchmark for evaluating language models across diverse subjects spanning humanities, social sciences, STEM, and professional knowledge.
Explore more about Benchmarks
Related terms
A more challenging version of MMLU designed to evaluate advanced reasoning and knowledge across diverse academic and professional subjects.
GPQAGPQAA benchmark that evaluates the ability of AI models to answer challenging graduate-level multiple-choice questions across scientific domains, measuring expert-level reasoning, scientific knowledge, and problem-solving performance.
MATHMATHA benchmark for evaluating mathematical problem-solving abilities of AI models across challenging competition-level mathematics problems.
Benchmark RunA single execution of a benchmark used to measure and record a model or system's performance under defined conditions.