Benchmark
AI2D
vision
reasoning
multimodal
multimodal
AI2D is a dataset of 4,903 illustrative diagrams from grade school natural sciences (such as food webs, human physiology, and life cycles) with over 15,000 multiple choice questions and answers. The benchmark evaluates diagram understanding and visual reasoning capabilities, requiring models to interpret diagrammatic elements, relationships, and structure to answer questions about scientific concepts represented in visual form.
语言EN
满分1
参评模型17
模型排名
| 名次 | 模型 | 机构 | 分数 | 来源 |
|---|---|---|---|---|
| 1 | Claude 3.5 Sonnet | Anthropic | 94.7 | 来源 ↗ |
| 2 | GPT-4o | OpenAI | 94.2 | 来源 ↗ |
| 3 | Pixtral Large | Mistral AI | 93.8 | 来源 ↗ |
| 4 | Mistral Small 3.2 24B Instruct | Mistral AI | 92.9 | 来源 ↗ |
| 5 | Llama 3.2 90B Instruct | Meta | 92.3 | 来源 ↗ |
| 6 | Llama 3.2 11B Instruct | Meta | 91.1 | 来源 ↗ |
| 7 | Qwen2.5 VL 72B Instruct | Alibaba Cloud / Qwen Team | 88.4 | 来源 ↗ |
| 8 | Grok-1.5V | xAI | 88.3 | 来源 ↗ |
| 9 | Gemma 3 27B | 84.5 | 来源 ↗ | |
| 10 | Gemma 3 12B | 84.2 | 来源 ↗ | |
| 11 | Qwen2.5-Omni-7B | Alibaba Cloud / Qwen Team | 83.2 | 来源 ↗ |
| 12 | Phi-4-multimodal-instruct | Microsoft | 82.3 | 来源 ↗ |
| 13 | DeepSeek VL2 | DeepSeek | 81.4 | 来源 ↗ |
| 14 | DeepSeek VL2 Small | DeepSeek | 80.0 | 来源 ↗ |
| 15 | Phi-3.5-vision-instruct | Microsoft | 78.1 | 来源 ↗ |
| 16 | Gemma 3 4B | 74.8 | 来源 ↗ | |
| 17 | DeepSeek VL2 Tiny | DeepSeek | 71.6 | 来源 ↗ |