MT-Bench
MT-BenchAlso called: Multi-Turn Benchmark
A benchmark for evaluating conversational AI models through multi-turn dialogue tasks that assess instruction following and response quality.
Explore more about Benchmarks
Related terms
A benchmark that evaluates advanced language models using difficult, real-world prompts and pairwise comparisons to measure instruction-following and overall response quality.
Chatbot ArenaA crowdsourced benchmark where users compare responses from multiple AI models through blind pairwise voting to rank model performance.
LLM-as-a-JudgeAn evaluation method that uses a large language model to assess the quality, correctness, or other attributes of AI-generated outputs.
Model ComparisonAn evaluation method that compares AI models on the same tasks, datasets, or metrics to identify differences in capabilities and performance.