Benchmark

VideoMMMU

multimodal vision reasoning multimodal

Video-MMMU evaluates Large Multimodal Models' ability to acquire knowledge from expert-level professional videos across six disciplines through three cognitive stages: perception, comprehension, and adaptation. Contains 300 videos and 900 human-annotated questions spanning Art, Business, Science, Medicine, Humanities, and Engineering.

语言EN
满分1
参评模型4

模型排名

名次 模型 机构 分数 来源
1 GPT-5 OpenAI 84.6 来源 ↗
2 Gemini 2.5 Pro Preview 06-05 Google 83.6 来源 ↗
3 o3 OpenAI 83.3 来源 ↗
4 GPT-4o OpenAI 61.2 来源 ↗