Benchmark

CharXiv-R

reasoning vision multimodal multimodal

CharXiv-R is the reasoning component of the CharXiv benchmark, focusing on complex reasoning questions that require synthesizing information across visual chart elements. It evaluates multimodal large language models on their ability to understand and reason about scientific charts from arXiv papers through various reasoning tasks.

语言EN
满分1
参评模型8

模型排名

名次 模型 机构 分数 来源
1 GPT-5 OpenAI 81.1 来源 ↗
2 o3 OpenAI 78.6 来源 ↗
3 o4-mini OpenAI 72.0 来源 ↗
4 GPT-4o OpenAI 58.8 来源 ↗
5 GPT-4.1 mini OpenAI 56.8 来源 ↗
6 GPT-4.1 OpenAI 56.7 来源 ↗
7 GPT-4.5 OpenAI 55.4 来源 ↗
8 GPT-4.1 nano OpenAI 40.5 来源 ↗