Benchmark

MMLU-Base

language reasoning math general text

Base version of the Massive Multitask Language Understanding benchmark, evaluating language models across 57 tasks including elementary mathematics, US history, computer science, law, and other professional and academic subjects. Designed to comprehensively measure the breadth and depth of a model's academic and professional understanding.

语言EN
满分1
参评模型1

模型排名

名次 模型 机构 分数 来源
1 Qwen2.5-Coder 7B Instruct Alibaba Cloud / Qwen Team 68.0 来源 ↗