Benchmark

FlenQA

reasoning long_context text

Flexible Length Question Answering dataset for evaluating the impact of input length on reasoning performance of language models, featuring True/False questions embedded in contexts of varying lengths (250-3000 tokens) across three reasoning tasks: Monotone Relations, People In Rooms, and simplified Ruletaker

语言EN
满分1
参评模型2

模型排名

名次 模型 机构 分数 来源
1 Phi 4 Reasoning Plus Microsoft 97.9 来源 ↗
2 Phi 4 Reasoning Microsoft 97.7 来源 ↗