Benchmark
MMMU (validation)
vision
multimodal
reasoning
general
multimodal
Validation set of the Massive Multi-discipline Multimodal Understanding and Reasoning benchmark. Features college-level multimodal questions across 6 core disciplines (Art & Design, Business, Science, Health & Medicine, Humanities & Social Science, Tech & Engineering) spanning 30 subjects and 183 subfields with diverse image types including charts, diagrams, maps, and tables.
语言EN
满分1
参评模型3
模型排名
| 名次 | 模型 | 机构 | 分数 | 来源 |
|---|---|---|---|---|
| 1 | Claude Opus 4.1 | Anthropic | 77.1 | 来源 ↗ |
| 2 | Claude Opus 4 | Anthropic | 76.5 | 来源 ↗ |
| 3 | Claude Haiku 4.5 | Anthropic | 73.2 | 来源 ↗ |