Benchmark
BoolQ
language
reasoning
text
BoolQ is a reading comprehension dataset for yes/no questions containing 15,942 naturally occurring examples. Each example consists of a question, passage, and boolean answer, where questions are generated in unprompted and unconstrained settings. The dataset challenges models with complex, non-factoid information requiring entailment-like inference to solve.
语言EN
满分1
参评模型9
模型排名
| 名次 | 模型 | 机构 | 分数 | 来源 |
|---|---|---|---|---|
| 1 | Gemma 2 27B | 84.8 | 来源 ↗ | |
| 2 | Phi-3.5-MoE-instruct | Microsoft | 84.6 | 来源 ↗ |
| 3 | Gemma 2 9B | 84.2 | 来源 ↗ | |
| 4 | Gemma 3n E4B | 81.6 | 来源 ↗ | |
| 5 | Gemma 3n E4B Instructed LiteRT Preview | 81.6 | 来源 ↗ | |
| 6 | Phi 4 Mini | Microsoft | 81.2 | 来源 ↗ |
| 7 | Phi-3.5-mini-instruct | Microsoft | 78.0 | 来源 ↗ |
| 8 | Gemma 3n E2B | 76.4 | 来源 ↗ | |
| 9 | Gemma 3n E2B Instructed LiteRT (Preview) | 76.4 | 来源 ↗ |