Benchmark

MMLU (CoT)

language reasoning math general text

Chain-of-Thought variant of the Massive Multitask Language Understanding benchmark, evaluating language models across 57 tasks including elementary mathematics, US history, computer science, law, and other professional and academic subjects. This version uses chain-of-thought prompting to elicit step-by-step reasoning.

语言EN
满分1
参评模型3

模型排名

名次 模型 机构 分数 来源
1 Llama 3.1 405B Instruct Meta 88.6 来源 ↗
2 Llama 3.1 70B Instruct Meta 86.0 来源 ↗
3 Llama 3.1 8B Instruct Meta 73.0 来源 ↗