Benchmark

CharXiv-D

reasoning vision multimodal multimodal

CharXiv-D is the descriptive questions subset of the CharXiv benchmark, designed to assess multimodal large language models' ability to extract basic information from scientific charts. It contains descriptive questions covering information extraction, enumeration, pattern recognition, and counting across 2,323 diverse charts from arXiv papers, all curated and verified by human experts.

语言EN
满分1
参评模型5

模型排名

名次 模型 机构 分数 来源
1 GPT-4.5 OpenAI 90.0 来源 ↗
2 GPT-4.1 mini OpenAI 88.4 来源 ↗
3 GPT-4.1 OpenAI 87.9 来源 ↗
4 GPT-4o OpenAI 85.3 来源 ↗
5 GPT-4.1 nano OpenAI 73.9 来源 ↗