Benchmark

ChartQA

reasoning vision multimodal multimodal

ChartQA is a large-scale benchmark comprising 9.6K human-written questions and 23.1K questions generated from human-written chart summaries, designed to evaluate models' abilities in visual and logical reasoning over charts.

语言EN
满分1
参评模型24

模型排名

名次 模型 机构 分数 来源
1 Claude 3.5 Sonnet Anthropic 90.8 来源 ↗
2 Llama 4 Maverick Meta 90.0 来源 ↗
3 Qwen2.5 VL 72B Instruct Alibaba Cloud / Qwen Team 89.5 来源 ↗
4 Nova Pro Amazon 89.2 来源 ↗
5 Llama 4 Scout Meta 88.8 来源 ↗
6 Qwen2-VL-72B-Instruct Alibaba Cloud / Qwen Team 88.3 来源 ↗
7 Pixtral Large Mistral AI 88.1 来源 ↗
8 Mistral Small 3.2 24B Instruct Mistral AI 87.4 来源 ↗
9 Qwen2.5 VL 7B Instruct Alibaba Cloud / Qwen Team 87.3 来源 ↗
10 Nova Lite Amazon 86.8 来源 ↗
11 DeepSeek VL2 DeepSeek 86.0 来源 ↗
12 GPT-4o OpenAI 85.7 来源 ↗
13 Llama 3.2 90B Instruct Meta 85.5 来源 ↗
14 Qwen2.5-Omni-7B Alibaba Cloud / Qwen Team 85.3 来源 ↗
15 DeepSeek VL2 Small DeepSeek 84.5 来源 ↗
16 Llama 3.2 11B Instruct Meta 83.4 来源 ↗
17 Pixtral-12B Mistral AI 81.8 来源 ↗
18 Phi-3.5-vision-instruct Microsoft 81.8 来源 ↗
19 Phi-4-multimodal-instruct Microsoft 81.4 来源 ↗
20 DeepSeek VL2 Tiny DeepSeek 81.0 来源 ↗
21 Gemma 3 27B Google 78.0 来源 ↗
22 Grok-1.5V xAI 76.1 来源 ↗
23 Gemma 3 12B Google 75.7 来源 ↗
24 Gemma 3 4B Google 68.8 来源 ↗