Benchmark

BoolQ

language reasoning text

BoolQ is a reading comprehension dataset for yes/no questions containing 15,942 naturally occurring examples. Each example consists of a question, passage, and boolean answer, where questions are generated in unprompted and unconstrained settings. The dataset challenges models with complex, non-factoid information requiring entailment-like inference to solve.

语言EN
满分1
参评模型9

模型排名

名次 模型 机构 分数 来源
1 Gemma 2 27B Google 84.8 来源 ↗
2 Phi-3.5-MoE-instruct Microsoft 84.6 来源 ↗
3 Gemma 2 9B Google 84.2 来源 ↗
4 Gemma 3n E4B Google 81.6 来源 ↗
5 Gemma 3n E4B Instructed LiteRT Preview Google 81.6 来源 ↗
6 Phi 4 Mini Microsoft 81.2 来源 ↗
7 Phi-3.5-mini-instruct Microsoft 78.0 来源 ↗
8 Gemma 3n E2B Google 76.4 来源 ↗
9 Gemma 3n E2B Instructed LiteRT (Preview) Google 76.4 来源 ↗