HellaSwag
A benchmark that evaluates the commonsense reasoning and natural language understanding capabilities of AI models by requiring them to select the most plausible continuation of a given context from multiple candidate endings.
Explore more about Benchmarks
Related terms
A benchmark for evaluating language models across diverse subjects spanning humanities, social sciences, STEM, and professional knowledge.
GSM8KGSM8KA benchmark that evaluates the ability of AI models to solve grade school mathematical word problems requiring multi-step reasoning, arithmetic, and logical problem-solving.
TruthfulQAA benchmark measuring how truthfully language models answer questions and avoid mimicking human falsehoods.