Chatbot Arena
Also called: LMSYS Chatbot Arena
A crowdsourced benchmark where users compare responses from multiple AI models through blind pairwise voting to rank model performance.
Explore more about Benchmarks
Related terms
A benchmark that evaluates advanced language models using difficult, real-world prompts and pairwise comparisons to measure instruction-following and overall response quality.
MT-BenchMT-BenchA benchmark for evaluating conversational AI models through multi-turn dialogue tasks that assess instruction following and response quality.
Pairwise ComparisonAn evaluation method that compares two outputs directly to determine which performs better against a defined criterion or preference.
Human EvaluationAn evaluation method in which human reviewers assess the quality of AI system outputs against defined criteria such as correctness, relevance, helpfulness, safety, or fluency, providing judgments that complement or validate automated evaluation metrics.
Blind EvaluationAn evaluation method where evaluators assess outputs without knowing which model, system, or approach produced them to reduce bias.