Benchmark
CharXiv-R
reasoning
vision
multimodal
multimodal
CharXiv-R is the reasoning component of the CharXiv benchmark, focusing on complex reasoning questions that require synthesizing information across visual chart elements. It evaluates multimodal large language models on their ability to understand and reason about scientific charts from arXiv papers through various reasoning tasks.
语言EN
满分1
参评模型8