Benchmark

InfoVQA

vision multimodal multimodal

InfoVQA dataset with 30,000 questions and 5,000 infographic images requiring joint reasoning over document layout, textual content, graphical elements, and data visualizations with elementary reasoning and arithmetic skills

语言EN
满分1
参评模型9

模型排名

名次 模型 机构 分数 来源
1 Qwen2.5 VL 32B Instruct Alibaba Cloud / Qwen Team 83.4 来源 ↗
2 Qwen2.5 VL 7B Instruct Alibaba Cloud / Qwen Team 82.6 来源 ↗
3 DeepSeek VL2 DeepSeek 78.1 来源 ↗
4 DeepSeek VL2 Small DeepSeek 75.8 来源 ↗
5 Phi-4-multimodal-instruct Microsoft 72.7 来源 ↗
6 Gemma 3 27B Google 70.6 来源 ↗
7 DeepSeek VL2 Tiny DeepSeek 66.1 来源 ↗
8 Gemma 3 12B Google 64.9 来源 ↗
9 Gemma 3 4B Google 50.0 来源 ↗