Benchmark
DocVQA
vision
multimodal
multimodal
A dataset for Visual Question Answering on document images containing 50,000 questions defined on 12,000+ document images. The benchmark tests AI's ability to understand document structure and content, requiring models to comprehend document layout and perform information retrieval to answer questions about document images.
语言EN
满分1
参评模型26
模型排名
| 名次 | 模型 | 机构 | 分数 | 来源 |
|---|---|---|---|---|
| 1 | Qwen2.5 VL 72B Instruct | Alibaba Cloud / Qwen Team | 96.4 | 来源 ↗ |
| 2 | Qwen2.5 VL 7B Instruct | Alibaba Cloud / Qwen Team | 95.7 | 来源 ↗ |
| 3 | Qwen2.5-Omni-7B | Alibaba Cloud / Qwen Team | 95.2 | 来源 ↗ |
| 4 | Claude 3.5 Sonnet | Anthropic | 95.2 | 来源 ↗ |
| 5 | Mistral Small 3.2 24B Instruct | Mistral AI | 94.9 | 来源 ↗ |
| 6 | Qwen2.5 VL 32B Instruct | Alibaba Cloud / Qwen Team | 94.8 | 来源 ↗ |
| 7 | Llama 4 Scout | Meta | 94.4 | 来源 ↗ |
| 8 | Llama 4 Maverick | Meta | 94.4 | 来源 ↗ |
| 9 | Grok-2 | xAI | 93.6 | 来源 ↗ |
| 10 | Nova Pro | Amazon | 93.5 | 来源 ↗ |
| 11 | Pixtral Large | Mistral AI | 93.3 | 来源 ↗ |
| 12 | DeepSeek VL2 | DeepSeek | 93.3 | 来源 ↗ |
| 13 | Grok-2 mini | xAI | 93.2 | 来源 ↗ |
| 14 | Phi-4-multimodal-instruct | Microsoft | 93.2 | 来源 ↗ |
| 15 | GPT-4o | OpenAI | 92.8 | 来源 ↗ |
| 16 | Nova Lite | Amazon | 92.4 | 来源 ↗ |
| 17 | DeepSeek VL2 Small | DeepSeek | 92.3 | 来源 ↗ |
| 18 | Pixtral-12B | Mistral AI | 90.7 | 来源 ↗ |
| 19 | Llama 3.2 90B Instruct | Meta | 90.1 | 来源 ↗ |
| 20 | DeepSeek VL2 Tiny | DeepSeek | 88.9 | 来源 ↗ |
| 21 | Llama 3.2 11B Instruct | Meta | 88.4 | 来源 ↗ |
| 22 | Gemma 3 12B | 87.1 | 来源 ↗ | |
| 23 | Gemma 3 27B | 86.6 | 来源 ↗ | |
| 24 | Grok-1.5V | xAI | 85.6 | 来源 ↗ |
| 25 | Grok-1.5 | xAI | 85.6 | 来源 ↗ |
| 26 | Gemma 3 4B | 75.8 | 来源 ↗ |