Benchmark

DocVQA

vision multimodal multimodal

A dataset for Visual Question Answering on document images containing 50,000 questions defined on 12,000+ document images. The benchmark tests AI's ability to understand document structure and content, requiring models to comprehend document layout and perform information retrieval to answer questions about document images.

语言EN
满分1
参评模型26

模型排名

名次 模型 机构 分数 来源
1 Qwen2.5 VL 72B Instruct Alibaba Cloud / Qwen Team 96.4 来源 ↗
2 Qwen2.5 VL 7B Instruct Alibaba Cloud / Qwen Team 95.7 来源 ↗
3 Qwen2.5-Omni-7B Alibaba Cloud / Qwen Team 95.2 来源 ↗
4 Claude 3.5 Sonnet Anthropic 95.2 来源 ↗
5 Mistral Small 3.2 24B Instruct Mistral AI 94.9 来源 ↗
6 Qwen2.5 VL 32B Instruct Alibaba Cloud / Qwen Team 94.8 来源 ↗
7 Llama 4 Scout Meta 94.4 来源 ↗
8 Llama 4 Maverick Meta 94.4 来源 ↗
9 Grok-2 xAI 93.6 来源 ↗
10 Nova Pro Amazon 93.5 来源 ↗
11 Pixtral Large Mistral AI 93.3 来源 ↗
12 DeepSeek VL2 DeepSeek 93.3 来源 ↗
13 Grok-2 mini xAI 93.2 来源 ↗
14 Phi-4-multimodal-instruct Microsoft 93.2 来源 ↗
15 GPT-4o OpenAI 92.8 来源 ↗
16 Nova Lite Amazon 92.4 来源 ↗
17 DeepSeek VL2 Small DeepSeek 92.3 来源 ↗
18 Pixtral-12B Mistral AI 90.7 来源 ↗
19 Llama 3.2 90B Instruct Meta 90.1 来源 ↗
20 DeepSeek VL2 Tiny DeepSeek 88.9 来源 ↗
21 Llama 3.2 11B Instruct Meta 88.4 来源 ↗
22 Gemma 3 12B Google 87.1 来源 ↗
23 Gemma 3 27B Google 86.6 来源 ↗
24 Grok-1.5V xAI 85.6 来源 ↗
25 Grok-1.5 xAI 85.6 来源 ↗
26 Gemma 3 4B Google 75.8 来源 ↗