Benchmark

Natural Questions

reasoning general search text

Natural Questions is a question answering dataset featuring real anonymized queries issued to Google search engine. It contains 307,373 training examples where annotators provide long answers (passages) and short answers (entities) from Wikipedia pages, or mark them as unanswerable.

语言EN
满分1
参评模型7

模型排名

名次 模型 机构 分数 来源
1 Gemma 2 27B Google 34.5 来源 ↗
2 Mistral NeMo Instruct Mistral AI 31.2 来源 ↗
3 Gemma 2 9B Google 29.2 来源 ↗
4 Gemma 3n E4B Google 20.9 来源 ↗
5 Gemma 3n E4B Instructed LiteRT Preview Google 20.9 来源 ↗
6 Gemma 3n E2B Google 15.5 来源 ↗
7 Gemma 3n E2B Instructed LiteRT (Preview) Google 15.5 来源 ↗