Benchmark
OpenAI MMLU
general
reasoning
math
legal
healthcare
finance
physics
chemistry
economics
psychology
text
MMLU (Massive Multitask Language Understanding) is a comprehensive benchmark that measures a text model's multitask accuracy across 57 diverse academic and professional subjects. The test covers elementary mathematics, US history, computer science, law, morality, business ethics, clinical knowledge, and many other domains spanning STEM, humanities, social sciences, and professional fields. To attain high accuracy, models must possess extensive world knowledge and problem-solving ability.
语言EN
满分1
参评模型2
模型排名
| 名次 | 模型 | 机构 | 分数 | 来源 |
|---|---|---|---|---|
| 1 | Gemma 3n E4B Instructed | 35.6 | 来源 ↗ | |
| 2 | Gemma 3n E2B Instructed | 22.3 | 来源 ↗ |