Benchmark

AI2 Reasoning Challenge (ARC)

reasoning general text

A dataset of 7,787 genuine grade-school level, multiple-choice science questions assembled to encourage research in advanced question-answering. The dataset is partitioned into a Challenge Set and Easy Set, where the Challenge Set contains only questions answered incorrectly by both retrieval-based and word co-occurrence algorithms. Covers multiple scientific domains including biology, physics, earth science, and chemistry, requiring scientific reasoning, causal understanding, and conceptual knowledge beyond simple fact retrieval. Includes a supporting corpus of over 14 million science sentences.

语言EN
满分1
参评模型1

模型排名

名次 模型 机构 分数 来源
1 GPT-4 OpenAI 96.3 来源 ↗