Benchmark

OpenBookQA

reasoning general text

OpenBookQA is a question-answering dataset modeled after open book exams for assessing human understanding. It contains 5,957 multiple-choice elementary-level science questions that probe understanding of 1,326 core science facts and their application to novel situations, requiring combination of open book facts with broad common knowledge through multi-hop reasoning.

语言EN
满分1
参评模型4

模型排名

名次 模型 机构 分数 来源
1 Phi-3.5-MoE-instruct Microsoft 89.6 来源 ↗
2 Phi-3.5-mini-instruct Microsoft 79.2 来源 ↗
3 Phi 4 Mini Microsoft 79.2 来源 ↗
4 Mistral NeMo Instruct Mistral AI 60.6 来源 ↗