Benchmark

MATH (CoT)

math reasoning text

MATH dataset contains 12,500 challenging competition mathematics problems from AMC 10, AMC 12, AIME, and other mathematics competitions. Each problem includes full step-by-step solutions and spans multiple difficulty levels (1-5) across seven mathematical subjects. This variant uses Chain-of-Thought prompting to encourage step-by-step reasoning.

语言EN
满分1
参评模型2

模型排名

名次 模型 机构 分数 来源
1 Llama 3.1 70B Instruct Meta 68.0 来源 ↗
2 Llama 3.1 8B Instruct Meta 51.9 来源 ↗