Benchmark

MMT-Bench

vision multimodal reasoning general multimodal

MMT-Bench is a comprehensive multimodal benchmark for evaluating Large Vision-Language Models towards multitask AGI. It comprises 31,325 meticulously curated multi-choice visual questions from various multimodal scenarios such as vehicle driving and embodied navigation, covering 32 core meta-tasks and 162 subtasks in multimodal understanding.

语言EN
满分1
参评模型4

模型排名

名次 模型 机构 分数 来源
1 DeepSeek VL2 DeepSeek 63.6 来源 ↗
2 Qwen2.5 VL 7B Instruct Alibaba Cloud / Qwen Team 63.6 来源 ↗
3 DeepSeek VL2 Small DeepSeek 62.9 来源 ↗
4 DeepSeek VL2 Tiny DeepSeek 53.2 来源 ↗