Benchmark

Multilingual MMLU

general reasoning language text 多语言

MMLU-ProX is a comprehensive multilingual benchmark covering 29 typologically diverse languages, building upon MMLU-Pro. Each language version consists of 11,829 identical questions enabling direct cross-linguistic comparisons. The benchmark evaluates large language models' reasoning capabilities across linguistic and cultural boundaries through challenging, reasoning-focused questions with 10 answer choices.

语言EN
满分1
参评模型2

模型排名

名次 模型 机构 分数 来源
1 o3-mini OpenAI 80.7 来源 ↗
2 Phi 4 Mini Microsoft 49.3 来源 ↗