Benchmark

MATH-500

math reasoning text

MATH-500 is a subset of the MATH dataset containing 500 challenging competition mathematics problems from AMC 10, AMC 12, AIME, and other mathematics competitions. Each problem includes full step-by-step solutions and spans multiple difficulty levels across seven mathematical subjects including Prealgebra, Algebra, Number Theory, Counting and Probability, Geometry, Intermediate Algebra, and Precalculus.

语言EN
满分1
参评模型25

模型排名

名次 模型 机构 分数 来源
1 GLM-4.5 Zhipu AI 98.2 来源 ↗
2 GLM-4.5-Air Zhipu AI 98.1 来源 ↗
3 Nemotron Nano 9B v2 NVIDIA 97.8 来源 ↗
4 Kimi K2-Instruct-0905 Moonshot AI 97.4 来源 ↗
5 Kimi K2 Instruct Moonshot AI 97.4 来源 ↗
6 Llama 3.1 Nemotron Ultra 253B v1 NVIDIA 97.0 来源 ↗
7 Llama-3.3 Nemotron Super 49B v1 NVIDIA 96.6 来源 ↗
8 Kimi-k1.5 Moonshot AI 96.2 来源 ↗
9 Claude 3.7 Sonnet Anthropic 96.2 来源 ↗
10 DeepSeek R1 Zero DeepSeek 95.9 来源 ↗
11 Llama 3.1 Nemotron Nano 8B V1 NVIDIA 95.4 来源 ↗
12 Phi 4 Mini Reasoning Microsoft 94.6 来源 ↗
13 DeepSeek R1 Distill Llama 70B DeepSeek 94.5 来源 ↗
14 DeepSeek R1 Distill Qwen 32B DeepSeek 94.3 来源 ↗
15 DeepSeek-V3 0324 DeepSeek 94.0 来源 ↗
16 DeepSeek R1 Distill Qwen 14B DeepSeek 93.9 来源 ↗
17 DeepSeek R1 Distill Qwen 7B DeepSeek 92.8 来源 ↗
18 QwQ-32B-Preview Alibaba Cloud / Qwen Team 90.6 来源 ↗
19 QwQ-32B Alibaba Cloud / Qwen Team 90.6 来源 ↗
20 DeepSeek-V3 DeepSeek 90.2 来源 ↗
21 o1-mini OpenAI 90.0 来源 ↗
22 DeepSeek R1 Distill Llama 8B DeepSeek 89.1 来源 ↗
23 DeepSeek R1 Distill Qwen 1.5B DeepSeek 83.9 来源 ↗
24 Granite 3.3 8B Base IBM 69.0 来源 ↗
25 Granite 3.3 8B Instruct IBM 69.0 来源 ↗