Benchmark

Video-MME

multimodal vision reasoning multimodal 多语言

Video-MME is the first-ever comprehensive evaluation benchmark of Multi-modal Large Language Models (MLLMs) in video analysis. It features 900 videos totaling 254 hours with 2,700 human-annotated question-answer pairs across 6 primary visual domains (Knowledge, Film & Television, Sports Competition, Life Record, Multilingual, and others) and 30 subfields. The benchmark evaluates models across diverse temporal dimensions (11 seconds to 1 hour), integrates multi-modal inputs including video frames, subtitles, and audio, and uses rigorous manual labeling by expert annotators for precise assessment.

语言EN
满分1
参评模型5

模型排名

名次 模型 机构 分数 来源
1 Gemini 2.5 Pro Google 84.8 来源 ↗
2 Gemini 1.5 Pro Google 78.6 来源 ↗
3 Gemini 1.5 Flash Google 76.1 来源 ↗
4 Gemini 1.5 Flash 8B Google 66.2 来源 ↗
5 Phi-4-multimodal-instruct Microsoft 55.0 来源 ↗