GPQA
GPQAAlso called: Graduate-Level Google-Proof Q&A
A benchmark that evaluates the ability of AI models to answer challenging graduate-level multiple-choice questions across scientific domains, measuring expert-level reasoning, scientific knowledge, and problem-solving performance.
Explore more about Benchmarks
Related terms
A benchmark for evaluating language models across diverse subjects spanning humanities, social sciences, STEM, and professional knowledge.
MMLU-ProA more challenging version of MMLU designed to evaluate advanced reasoning and knowledge across diverse academic and professional subjects.
TruthfulQAA benchmark measuring how truthfully language models answer questions and avoid mimicking human falsehoods.
HumanEvalA benchmark that evaluates the code generation capabilities of AI models by measuring their ability to generate functionally correct code that passes predefined unit tests for a collection of programming tasks.
Reference-based EvaluationAn evaluation method that measures model output quality by comparing generated responses against ground-truth reference data or golden datasets.