Arena-Hard
Also called: Arena Hard
A benchmark that evaluates advanced language models using difficult, real-world prompts and pairwise comparisons to measure instruction-following and overall response quality.
Explore more about Benchmarks
Related terms
A crowdsourced benchmark where users compare responses from multiple AI models through blind pairwise voting to rank model performance.
MT-BenchMT-BenchA benchmark for evaluating conversational AI models through multi-turn dialogue tasks that assess instruction following and response quality.
MMLUMMLUA benchmark for evaluating language models across diverse subjects spanning humanities, social sciences, STEM, and professional knowledge.
LLM-as-a-JudgeAn evaluation method that uses a large language model to assess the quality, correctness, or other attributes of AI-generated outputs.
Pairwise ComparisonAn evaluation method that compares two outputs directly to determine which performs better against a defined criterion or preference.