Benchmark

MATH

math reasoning text

MATH dataset contains 12,500 challenging competition mathematics problems from AMC 10, AMC 12, AIME, and other mathematics competitions. Each problem includes full step-by-step solutions and spans multiple difficulty levels (1-5) across seven mathematical subjects including Prealgebra, Algebra, Number Theory, Counting and Probability, Geometry, Intermediate Algebra, and Precalculus.

语言EN
满分1
参评模型64

模型排名

名次 模型 机构 分数 来源
1 o3-mini OpenAI 97.9 来源 ↗
2 o1 OpenAI 96.4 来源 ↗
3 Gemini 2.0 Flash Google 89.7 来源 ↗
4 Kimi K2 0905 Moonshot AI 89.1 来源 ↗
5 Gemma 3 27B Google 89.0 来源 ↗
6 Gemini 2.0 Flash-Lite Google 86.8 来源 ↗
7 Gemini 1.5 Pro Google 86.5 来源 ↗
8 o1-preview OpenAI 85.5 来源 ↗
9 GPT-5 OpenAI 84.7 来源 ↗
10 Gemma 3 12B Google 83.8 来源 ↗
11 Qwen2.5 72B Instruct Alibaba Cloud / Qwen Team 83.1 来源 ↗
12 Qwen2.5 32B Instruct Alibaba Cloud / Qwen Team 83.1 来源 ↗
13 Qwen2.5 VL 32B Instruct Alibaba Cloud / Qwen Team 82.2 来源 ↗
14 Phi 4 Microsoft 80.4 来源 ↗
15 Qwen2.5 14B Instruct Alibaba Cloud / Qwen Team 80.0 来源 ↗
16 Claude 3.5 Sonnet Anthropic 78.3 来源 ↗
17 Gemini 1.5 Flash Google 77.9 来源 ↗
18 Llama 3.3 70B Instruct Meta 77.0 来源 ↗
19 GPT-4o OpenAI 76.6 来源 ↗
20 Nova Pro Amazon 76.6 来源 ↗
21 Grok-2 xAI 76.1 来源 ↗
22 Gemma 3 4B Google 75.6 来源 ↗
23 Qwen2.5 7B Instruct Alibaba Cloud / Qwen Team 75.5 来源 ↗
24 DeepSeek-V2.5 DeepSeek 74.7 来源 ↗
25 Llama 3.1 405B Instruct Meta 73.8 来源 ↗
26 Nova Lite Amazon 73.3 来源 ↗
27 Grok-2 mini xAI 73.0 来源 ↗
28 GPT-4 Turbo OpenAI 72.6 来源 ↗
29 Qwen3 235B A22B Alibaba Cloud / Qwen Team 71.8 来源 ↗
30 Qwen2.5-Omni-7B Alibaba Cloud / Qwen Team 71.5 来源 ↗
31 Claude 3.5 Sonnet Anthropic 71.1 来源 ↗
32 Mistral Small 3 24B Instruct Mistral AI 70.6 来源 ↗
33 GPT-4o mini OpenAI 70.2 来源 ↗
34 Kimi K2 Base Moonshot AI 70.2 来源 ↗
35 Mistral Small 3.2 24B Instruct Mistral AI 69.4 来源 ↗
36 Claude 3.5 Haiku Anthropic 69.4 来源 ↗
37 Mistral Small 3.1 24B Instruct Mistral AI 69.3 来源 ↗
38 Nova Micro Amazon 69.3 来源 ↗
39 Llama 3.2 90B Instruct Meta 68.0 来源 ↗
40 Phi 4 Mini Microsoft 64.0 来源 ↗
41 Llama 4 Maverick Meta 61.2 来源 ↗
42 Claude 3 Opus Anthropic 60.1 来源 ↗
43 Qwen2 72B Instruct Alibaba Cloud / Qwen Team 59.7 来源 ↗
44 Phi-3.5-MoE-instruct Microsoft 59.5 来源 ↗
45 Gemini 1.5 Flash 8B Google 58.7 来源 ↗
46 Qwen2.5-Coder 32B Instruct Alibaba Cloud / Qwen Team 57.2 来源 ↗
47 Ministral 8B Instruct Mistral AI 54.5 来源 ↗
48 Llama 3.2 11B Instruct Meta 51.9 来源 ↗
49 Grok-1.5 xAI 50.6 来源 ↗
50 Llama 4 Scout Meta 50.3 来源 ↗
51 Qwen2 7B Instruct Alibaba Cloud / Qwen Team 49.6 来源 ↗
52 Phi-3.5-mini-instruct Microsoft 48.5 来源 ↗
53 Pixtral-12B Mistral AI 48.1 来源 ↗
54 Llama 3.2 3B Instruct Meta 48.0 来源 ↗
55 Gemma 3 1B Google 48.0 来源 ↗
56 Qwen2.5-Coder 7B Instruct Alibaba Cloud / Qwen Team 46.6 来源 ↗
57 Mistral Small 3 24B Base Mistral AI 46.0 来源 ↗
58 GPT-3.5 Turbo OpenAI 43.1 来源 ↗
59 Claude 3 Sonnet Anthropic 43.1 来源 ↗
60 Gemma 2 27B Google 42.3 来源 ↗
61 GPT-4 OpenAI 42.0 来源 ↗
62 Claude 3 Haiku Anthropic 38.9 来源 ↗
63 Gemma 2 9B Google 36.6 来源 ↗
64 Gemini 1.0 Pro Google 32.6 来源 ↗