Benchmark
CharXiv-D
reasoning
vision
multimodal
multimodal
CharXiv-D is the descriptive questions subset of the CharXiv benchmark, designed to assess multimodal large language models' ability to extract basic information from scientific charts. It contains descriptive questions covering information extraction, enumeration, pattern recognition, and counting across 2,323 diverse charts from arXiv papers, all curated and verified by human experts.
语言EN
满分1
参评模型5
模型排名
| 名次 | 模型 | 机构 | 分数 | 来源 |
|---|---|---|---|---|
| 1 | GPT-4.5 | OpenAI | 90.0 | 来源 ↗ |
| 2 | GPT-4.1 mini | OpenAI | 88.4 | 来源 ↗ |
| 3 | GPT-4.1 | OpenAI | 87.9 | 来源 ↗ |
| 4 | GPT-4o | OpenAI | 85.3 | 来源 ↗ |
| 5 | GPT-4.1 nano | OpenAI | 73.9 | 来源 ↗ |