Benchmark
OpenBookQA
reasoning
general
text
OpenBookQA is a question-answering dataset modeled after open book exams for assessing human understanding. It contains 5,957 multiple-choice elementary-level science questions that probe understanding of 1,326 core science facts and their application to novel situations, requiring combination of open book facts with broad common knowledge through multi-hop reasoning.
语言EN
满分1
参评模型4
模型排名
| 名次 | 模型 | 机构 | 分数 | 来源 |
|---|---|---|---|---|
| 1 | Phi-3.5-MoE-instruct | Microsoft | 89.6 | 来源 ↗ |
| 2 | Phi-3.5-mini-instruct | Microsoft | 79.2 | 来源 ↗ |
| 3 | Phi 4 Mini | Microsoft | 79.2 | 来源 ↗ |
| 4 | Mistral NeMo Instruct | Mistral AI | 60.6 | 来源 ↗ |