MATH
MATHAlso called: MATH Benchmark
A benchmark for evaluating mathematical problem-solving abilities of AI models across challenging competition-level mathematics problems.
Explore more about Benchmarks
Related terms
A benchmark that evaluates the ability of AI models to solve grade school mathematical word problems requiring multi-step reasoning, arithmetic, and logical problem-solving.
MMLUMMLUA benchmark for evaluating language models across diverse subjects spanning humanities, social sciences, STEM, and professional knowledge.
MMLU-ProA more challenging version of MMLU designed to evaluate advanced reasoning and knowledge across diverse academic and professional subjects.
Benchmark RunA single execution of a benchmark used to measure and record a model or system's performance under defined conditions.