ARC
ARCAlso called: AI2 Reasoning Challenge
A benchmark that evaluates an AI model's ability to solve grade-school science questions requiring reasoning, knowledge, and commonsense understanding.
Explore more about Benchmarks
Related terms
A benchmark for evaluating language models across diverse subjects spanning humanities, social sciences, STEM, and professional knowledge.
GPQAGPQAA benchmark that evaluates the ability of AI models to answer challenging graduate-level multiple-choice questions across scientific domains, measuring expert-level reasoning, scientific knowledge, and problem-solving performance.
HellaSwagA benchmark that evaluates the commonsense reasoning and natural language understanding capabilities of AI models by requiring them to select the most plausible continuation of a given context from multiple candidate endings.
TruthfulQAA benchmark measuring how truthfully language models answer questions and avoid mimicking human falsehoods.
GSM8KGSM8KA benchmark that evaluates the ability of AI models to solve grade school mathematical word problems requiring multi-step reasoning, arithmetic, and logical problem-solving.