Benchmark

MMBench

vision multimodal reasoning multimodal 多语言

A bilingual benchmark for assessing multi-modal capabilities of vision-language models through multiple-choice questions in both English and Chinese, providing systematic evaluation across diverse vision-language tasks with robust metrics.

语言EN
满分1
参评模型7

模型排名

名次 模型 机构 分数 来源
1 Qwen2.5 VL 72B Instruct Alibaba Cloud / Qwen Team 88.0 来源 ↗
2 Phi-4-multimodal-instruct Microsoft 86.7 来源 ↗
3 Qwen2.5 VL 7B Instruct Alibaba Cloud / Qwen Team 84.3 来源 ↗
4 Phi-3.5-vision-instruct Microsoft 81.9 来源 ↗
5 DeepSeek VL2 Small DeepSeek 80.3 来源 ↗
6 DeepSeek VL2 DeepSeek 79.6 来源 ↗
7 DeepSeek VL2 Tiny DeepSeek 69.2 来源 ↗