LongBench
A benchmark for evaluating large language models on tasks requiring understanding and reasoning over long contexts.
Explore more about Benchmarks
Related terms
A benchmark for evaluating language models across diverse subjects spanning humanities, social sciences, STEM, and professional knowledge.
MMLU-ProA more challenging version of MMLU designed to evaluate advanced reasoning and knowledge across diverse academic and professional subjects.
MMMUMMMUA multimodal benchmark for evaluating models on complex tasks spanning diverse academic disciplines and requiring both visual and textual reasoning.
Benchmark RunA single execution of a benchmark used to measure and record a model or system's performance under defined conditions.