Benchmark

MMLU

general reasoning language math text

Massive Multitask Language Understanding benchmark testing knowledge across 57 diverse subjects including STEM, humanities, social sciences, and professional domains

语言EN
满分1
参评模型80

模型排名

名次 模型 机构 分数 来源
1 GPT-5 OpenAI 92.5 来源 ↗
2 o1 OpenAI 91.8 来源 ↗
3 o1-preview OpenAI 90.8 来源 ↗
4 GPT-4.5 OpenAI 90.8 来源 ↗
5 Claude 3.5 Sonnet Anthropic 90.4 来源 ↗
6 Claude 3.5 Sonnet Anthropic 90.4 来源 ↗
7 GPT-4.1 OpenAI 90.2 来源 ↗
8 Kimi K2 0905 Moonshot AI 90.2 来源 ↗
9 GPT OSS 120B OpenAI 90.0 来源 ↗
10 Kimi K2-Instruct-0905 Moonshot AI 89.5 来源 ↗
11 Kimi K2 Instruct Moonshot AI 89.5 来源 ↗
12 GPT-4o OpenAI 88.7 来源 ↗
13 DeepSeek-V3 DeepSeek 88.5 来源 ↗
14 Qwen3 235B A22B Alibaba Cloud / Qwen Team 87.8 来源 ↗
15 Kimi K2 Base Moonshot AI 87.8 来源 ↗
16 Grok-2 xAI 87.5 来源 ↗
17 GPT-4.1 mini OpenAI 87.5 来源 ↗
18 Kimi-k1.5 Moonshot AI 87.4 来源 ↗
19 Llama 3.1 405B Instruct Meta 87.3 来源 ↗
20 o3-mini OpenAI 86.9 来源 ↗
21 Claude 3 Opus Anthropic 86.8 来源 ↗
22 GPT-4 Turbo OpenAI 86.5 来源 ↗
23 GPT-4 OpenAI 86.4 来源 ↗
24 Grok-2 mini xAI 86.2 来源 ↗
25 Llama 3.3 70B Instruct Meta 86.0 来源 ↗
26 Llama 3.2 90B Instruct Meta 86.0 来源 ↗
27 Gemini 1.5 Pro Google 85.9 来源 ↗
28 Nova Pro Amazon 85.9 来源 ↗
29 GPT-4o OpenAI 85.7 来源 ↗
30 Llama 4 Maverick Meta 85.5 来源 ↗
31 GPT OSS 20B OpenAI 85.3 来源 ↗
32 o1-mini OpenAI 85.2 来源 ↗
33 Phi 4 Microsoft 84.8 来源 ↗
34 Mistral Large 2 Mistral AI 84.0 来源 ↗
35 Llama 3.1 70B Instruct Meta 83.6 来源 ↗
36 Qwen2.5 32B Instruct Alibaba Cloud / Qwen Team 83.3 来源 ↗
37 Qwen2 72B Instruct Alibaba Cloud / Qwen Team 82.3 来源 ↗
38 GPT-4o mini OpenAI 82.0 来源 ↗
39 Grok-1.5 xAI 81.3 来源 ↗
40 Jamba 1.5 Large AI21 Labs 81.2 来源 ↗
41 Mistral Small 3.1 24B Base Mistral AI 81.0 来源 ↗
42 Mistral Small 3 24B Base Mistral AI 80.7 来源 ↗
43 Mistral Small 3.1 24B Instruct Mistral AI 80.6 来源 ↗
44 Mistral Small 3.2 24B Instruct Mistral AI 80.5 来源 ↗
45 Nova Lite Amazon 80.5 来源 ↗
46 DeepSeek-V2.5 DeepSeek 80.4 来源 ↗
47 Llama 3.1 Nemotron 70B Instruct NVIDIA 80.2 来源 ↗
48 GPT-4.1 nano OpenAI 80.1 来源 ↗
49 Qwen2.5 14B Instruct Alibaba Cloud / Qwen Team 79.7 来源 ↗
50 Llama 4 Scout Meta 79.6 来源 ↗
51 Claude 3 Sonnet Anthropic 79.0 来源 ↗
52 Phi-3.5-MoE-instruct Microsoft 78.9 来源 ↗
53 Gemini 1.5 Flash Google 78.9 来源 ↗
54 Qwen2.5 VL 32B Instruct Alibaba Cloud / Qwen Team 78.4 来源 ↗
55 Nova Micro Amazon 77.6 来源 ↗
56 Command R+ Cohere 75.7 来源 ↗
57 Gemma 2 27B Google 75.2 来源 ↗
58 Claude 3 Haiku Anthropic 75.2 来源 ↗
59 Qwen2.5-Coder 32B Instruct Alibaba Cloud / Qwen Team 75.1 来源 ↗
60 Llama 3.2 11B Instruct Meta 73.0 来源 ↗
61 Gemini 1.0 Pro Google 71.8 来源 ↗
62 Gemma 2 9B Google 71.3 来源 ↗
63 Qwen2 7B Instruct Alibaba Cloud / Qwen Team 70.5 来源 ↗
64 GPT-3.5 Turbo OpenAI 69.8 来源 ↗
65 Jamba 1.5 Mini AI21 Labs 69.7 来源 ↗
66 Llama 3.1 8B Instruct Meta 69.4 来源 ↗
67 Pixtral-12B Mistral AI 69.2 来源 ↗
68 Phi-3.5-mini-instruct Microsoft 69.0 来源 ↗
69 Mistral NeMo Instruct Mistral AI 68.0 来源 ↗
70 Qwen2.5-Coder 7B Instruct Alibaba Cloud / Qwen Team 67.6 来源 ↗
71 Phi 4 Mini Microsoft 67.3 来源 ↗
72 Granite 3.3 8B Instruct IBM 65.5 来源 ↗
73 Ministral 8B Instruct Mistral AI 65.0 来源 ↗
74 Gemma 3n E4B Instructed LiteRT Preview Google 64.9 来源 ↗
75 Gemma 3n E4B Instructed Google 64.9 来源 ↗
76 Granite 3.3 8B Base IBM 63.9 来源 ↗
77 Llama 3.2 3B Instruct Meta 63.4 来源 ↗
78 IBM Granite 4.0 Tiny Preview IBM 60.4 来源 ↗
79 Gemma 3n E2B Instructed Google 60.1 来源 ↗
80 Gemma 3n E2B Instructed LiteRT (Preview) Google 60.1 来源 ↗