Benchmark
MVBench
vision
video
multimodal
spatial_reasoning
reasoning
multimodal
A comprehensive multi-modal video understanding benchmark covering 20 challenging video tasks that require temporal understanding beyond single-frame analysis. Tasks span from perception to cognition, including action recognition, temporal reasoning, spatial reasoning, object interaction, scene transition, and counterfactual inference. Uses a novel static-to-dynamic method to systematically generate video tasks from existing annotations.
语言EN
满分1
参评模型4
模型排名
| 名次 | 模型 | 机构 | 分数 | 来源 |
|---|---|---|---|---|
| 1 | Qwen2-VL-72B-Instruct | Alibaba Cloud / Qwen Team | 73.6 | 来源 ↗ |
| 2 | Qwen2.5 VL 72B Instruct | Alibaba Cloud / Qwen Team | 70.4 | 来源 ↗ |
| 3 | Qwen2.5-Omni-7B | Alibaba Cloud / Qwen Team | 70.3 | 来源 ↗ |
| 4 | Qwen2.5 VL 7B Instruct | Alibaba Cloud / Qwen Team | 69.6 | 来源 ↗ |