Benchmark

Qasper

reasoning long_context text

QASPER is a dataset of 5,049 information-seeking questions and answers anchored in 1,585 NLP research papers. Questions are written by NLP practitioners who read only titles and abstracts, while answers require understanding the full paper text and provide supporting evidence. The dataset challenges models with complex reasoning across document sections for academic document question answering. Each question seeks information present in the full text and is answered by a separate set of NLP practitioners who also provide supporting evidence to answers.

语言EN
满分1
参评模型2

模型排名

名次 模型 机构 分数 来源
1 Phi-3.5-mini-instruct Microsoft 41.9 来源 ↗
2 Phi-3.5-MoE-instruct Microsoft 40.0 来源 ↗