Benchmark

MMStar

vision multimodal reasoning general multimodal

MMStar is an elite vision-indispensable multimodal benchmark comprising 1,500 challenge samples meticulously selected by humans to evaluate 6 core capabilities and 18 detailed axes. The benchmark addresses issues of visual content unnecessity and unintentional data leakage in existing multimodal evaluations.

语言EN
满分1
参评模型7

模型排名

名次 模型 机构 分数 来源
1 Qwen2.5 VL 72B Instruct Alibaba Cloud / Qwen Team 70.8 来源 ↗
2 Qwen2.5 VL 32B Instruct Alibaba Cloud / Qwen Team 69.5 来源 ↗
3 Qwen2.5-Omni-7B Alibaba Cloud / Qwen Team 64.0 来源 ↗
4 Qwen2.5 VL 7B Instruct Alibaba Cloud / Qwen Team 63.9 来源 ↗
5 DeepSeek VL2 DeepSeek 61.3 来源 ↗
6 DeepSeek VL2 Small DeepSeek 57.0 来源 ↗
7 DeepSeek VL2 Tiny DeepSeek 45.9 来源 ↗