Benchmark

MME

vision multimodal reasoning multimodal

A comprehensive evaluation benchmark for Multimodal Large Language Models measuring both perception and cognition abilities across 14 subtasks. Features manually designed instruction-answer pairs to avoid data leakage and provides systematic quantitative assessment of MLLM capabilities.

语言EN
满分1
参评模型3

模型排名

名次 模型 机构 分数 来源
1 DeepSeek VL2 DeepSeek 22.5 来源 ↗
2 DeepSeek VL2 Small DeepSeek 21.2 来源 ↗
3 DeepSeek VL2 Tiny DeepSeek 19.2 来源 ↗