Benchmark

ARC-C

reasoning general text

The AI2 Reasoning Challenge (ARC) Challenge Set is a multiple-choice question-answering benchmark containing grade-school level science questions that require advanced reasoning capabilities. ARC-C specifically contains questions that were answered incorrectly by both retrieval-based and word co-occurrence algorithms, making it a particularly challenging subset designed to test commonsense reasoning abilities in AI systems.

语言EN
满分1
参评模型31

模型排名

名次 模型 机构 分数 来源
1 Llama 3.1 405B Instruct Meta 96.9 来源 ↗
2 Claude 3 Opus Anthropic 96.4 来源 ↗
3 Llama 3.1 70B Instruct Meta 94.8 来源 ↗
4 Nova Pro Amazon 94.8 来源 ↗
5 Claude 3 Sonnet Anthropic 93.2 来源 ↗
6 Jamba 1.5 Large AI21 Labs 93.0 来源 ↗
7 Nova Lite Amazon 92.4 来源 ↗
8 Mistral Small 3 24B Base Mistral AI 91.3 来源 ↗
9 Phi-3.5-MoE-instruct Microsoft 91.0 来源 ↗
10 Nova Micro Amazon 90.2 来源 ↗
11 Claude 3 Haiku Anthropic 89.2 来源 ↗
12 Jamba 1.5 Mini AI21 Labs 85.7 来源 ↗
13 Phi-3.5-mini-instruct Microsoft 84.6 来源 ↗
14 Phi 4 Mini Microsoft 83.7 来源 ↗
15 Llama 3.1 8B Instruct Meta 83.4 来源 ↗
16 Llama 3.2 3B Instruct Meta 78.6 来源 ↗
17 Ministral 8B Instruct Mistral AI 71.9 来源 ↗
18 Gemma 2 27B Google 71.4 来源 ↗
19 Command R+ Cohere 71.0 来源 ↗
20 Qwen2.5-Coder 32B Instruct Alibaba Cloud / Qwen Team 70.5 来源 ↗
21 Qwen2.5 32B Instruct Alibaba Cloud / Qwen Team 70.4 来源 ↗
22 Llama 3.1 Nemotron 70B Instruct NVIDIA 69.2 来源 ↗
23 Qwen2 72B Instruct Alibaba Cloud / Qwen Team 68.9 来源 ↗
24 Gemma 2 9B Google 68.4 来源 ↗
25 Qwen2.5 14B Instruct Alibaba Cloud / Qwen Team 67.3 来源 ↗
26 Gemma 3n E4B Instructed LiteRT Preview Google 61.6 来源 ↗
27 Gemma 3n E4B Google 61.6 来源 ↗
28 Qwen2.5-Coder 7B Instruct Alibaba Cloud / Qwen Team 60.9 来源 ↗
29 Gemma 3n E2B Instructed LiteRT (Preview) Google 51.7 来源 ↗
30 Gemma 3n E2B Google 51.7 来源 ↗
31 Granite 3.3 8B Base IBM 50.8 来源 ↗