Benchmark

AI2D

vision reasoning multimodal multimodal

AI2D is a dataset of 4,903 illustrative diagrams from grade school natural sciences (such as food webs, human physiology, and life cycles) with over 15,000 multiple choice questions and answers. The benchmark evaluates diagram understanding and visual reasoning capabilities, requiring models to interpret diagrammatic elements, relationships, and structure to answer questions about scientific concepts represented in visual form.

语言EN
满分1
参评模型17

模型排名

名次 模型 机构 分数 来源
1 Claude 3.5 Sonnet Anthropic 94.7 来源 ↗
2 GPT-4o OpenAI 94.2 来源 ↗
3 Pixtral Large Mistral AI 93.8 来源 ↗
4 Mistral Small 3.2 24B Instruct Mistral AI 92.9 来源 ↗
5 Llama 3.2 90B Instruct Meta 92.3 来源 ↗
6 Llama 3.2 11B Instruct Meta 91.1 来源 ↗
7 Qwen2.5 VL 72B Instruct Alibaba Cloud / Qwen Team 88.4 来源 ↗
8 Grok-1.5V xAI 88.3 来源 ↗
9 Gemma 3 27B Google 84.5 来源 ↗
10 Gemma 3 12B Google 84.2 来源 ↗
11 Qwen2.5-Omni-7B Alibaba Cloud / Qwen Team 83.2 来源 ↗
12 Phi-4-multimodal-instruct Microsoft 82.3 来源 ↗
13 DeepSeek VL2 DeepSeek 81.4 来源 ↗
14 DeepSeek VL2 Small DeepSeek 80.0 来源 ↗
15 Phi-3.5-vision-instruct Microsoft 78.1 来源 ↗
16 Gemma 3 4B Google 74.8 来源 ↗
17 DeepSeek VL2 Tiny DeepSeek 71.6 来源 ↗