Benchmark
MMStar
vision
multimodal
reasoning
general
multimodal
MMStar is an elite vision-indispensable multimodal benchmark comprising 1,500 challenge samples meticulously selected by humans to evaluate 6 core capabilities and 18 detailed axes. The benchmark addresses issues of visual content unnecessity and unintentional data leakage in existing multimodal evaluations.
语言EN
满分1
参评模型7
模型排名
| 名次 | 模型 | 机构 | 分数 | 来源 |
|---|---|---|---|---|
| 1 | Qwen2.5 VL 72B Instruct | Alibaba Cloud / Qwen Team | 70.8 | 来源 ↗ |
| 2 | Qwen2.5 VL 32B Instruct | Alibaba Cloud / Qwen Team | 69.5 | 来源 ↗ |
| 3 | Qwen2.5-Omni-7B | Alibaba Cloud / Qwen Team | 64.0 | 来源 ↗ |
| 4 | Qwen2.5 VL 7B Instruct | Alibaba Cloud / Qwen Team | 63.9 | 来源 ↗ |
| 5 | DeepSeek VL2 | DeepSeek | 61.3 | 来源 ↗ |
| 6 | DeepSeek VL2 Small | DeepSeek | 57.0 | 来源 ↗ |
| 7 | DeepSeek VL2 Tiny | DeepSeek | 45.9 | 来源 ↗ |