GSM8K
GSM8KAlso called: Grade School Math 8K
A benchmark that evaluates the ability of AI models to solve grade school mathematical word problems requiring multi-step reasoning, arithmetic, and logical problem-solving.
Explore more about Benchmarks
Related terms
A benchmark for evaluating mathematical problem-solving abilities of AI models across challenging competition-level mathematics problems.
MMLUMMLUA benchmark for evaluating language models across diverse subjects spanning humanities, social sciences, STEM, and professional knowledge.
GPQAGPQAA benchmark that evaluates the ability of AI models to answer challenging graduate-level multiple-choice questions across scientific domains, measuring expert-level reasoning, scientific knowledge, and problem-solving performance.
HumanEvalA benchmark that evaluates the code generation capabilities of AI models by measuring their ability to generate functionally correct code that passes predefined unit tests for a collection of programming tasks.
Chain of ThoughtCoTA reasoning technique in which an AI model generates intermediate reasoning steps before producing a final answer or decision.