Benchmark

MVBench

vision video multimodal spatial_reasoning reasoning multimodal

A comprehensive multi-modal video understanding benchmark covering 20 challenging video tasks that require temporal understanding beyond single-frame analysis. Tasks span from perception to cognition, including action recognition, temporal reasoning, spatial reasoning, object interaction, scene transition, and counterfactual inference. Uses a novel static-to-dynamic method to systematically generate video tasks from existing annotations.

语言EN
满分1
参评模型4

模型排名

名次 模型 机构 分数 来源
1 Qwen2-VL-72B-Instruct Alibaba Cloud / Qwen Team 73.6 来源 ↗
2 Qwen2.5 VL 72B Instruct Alibaba Cloud / Qwen Team 70.4 来源 ↗
3 Qwen2.5-Omni-7B Alibaba Cloud / Qwen Team 70.3 来源 ↗
4 Qwen2.5 VL 7B Instruct Alibaba Cloud / Qwen Team 69.6 来源 ↗