DocVQA
DocVQAAlso called: Document Visual Question Answering
A benchmark for evaluating an AI model's ability to answer questions by understanding and extracting information from document images.
Explore more about Benchmarks
Related terms
A multimodal benchmark for evaluating models on complex tasks spanning diverse academic disciplines and requiring both visual and textual reasoning.
MMBenchMMBenchA benchmark for evaluating multimodal large language models across a broad range of vision-language tasks and capabilities.
Human EvaluationAn evaluation method in which human reviewers assess the quality of AI system outputs against defined criteria such as correctness, relevance, helpfulness, safety, or fluency, providing judgments that complement or validate automated evaluation metrics.
Reference-based EvaluationAn evaluation method that measures model output quality by comparing generated responses against ground-truth reference data or golden datasets.