Benchmark
Video-MME
multimodal
vision
reasoning
multimodal
多语言
Video-MME is the first-ever comprehensive evaluation benchmark of Multi-modal Large Language Models (MLLMs) in video analysis. It features 900 videos totaling 254 hours with 2,700 human-annotated question-answer pairs across 6 primary visual domains (Knowledge, Film & Television, Sports Competition, Life Record, Multilingual, and others) and 30 subfields. The benchmark evaluates models across diverse temporal dimensions (11 seconds to 1 hour), integrates multi-modal inputs including video frames, subtitles, and audio, and uses rigorous manual labeling by expert annotators for precise assessment.
语言EN
满分1
参评模型5
模型排名
| 名次 | 模型 | 机构 | 分数 | 来源 |
|---|---|---|---|---|
| 1 | Gemini 2.5 Pro | 84.8 | 来源 ↗ | |
| 2 | Gemini 1.5 Pro | 78.6 | 来源 ↗ | |
| 3 | Gemini 1.5 Flash | 76.1 | 来源 ↗ | |
| 4 | Gemini 1.5 Flash 8B | 66.2 | 来源 ↗ | |
| 5 | Phi-4-multimodal-instruct | Microsoft | 55.0 | 来源 ↗ |