Benchmark

MMBench-V1.1

vision multimodal reasoning multimodal 多语言

Version 1.1 of MMBench, an improved bilingual benchmark for assessing multi-modal capabilities of vision-language models through multiple-choice questions in both English and Chinese, providing systematic evaluation across diverse vision-language tasks.

语言EN
满分1
参评模型4

模型排名

名次 模型 机构 分数 来源
1 Qwen2.5-Omni-7B Alibaba Cloud / Qwen Team 81.8 来源 ↗
2 DeepSeek VL2 Small DeepSeek 79.3 来源 ↗
3 DeepSeek VL2 DeepSeek 79.2 来源 ↗
4 DeepSeek VL2 Tiny DeepSeek 68.3 来源 ↗