Benchmark

MRCR

long_context reasoning general text

MRCR (Multi-Round Coreference Resolution) is a synthetic long-context reasoning task where models must navigate long conversations to reproduce specific model outputs. It tests the ability to distinguish between similar requests and reason about ordering while maintaining attention across extended contexts.

语言EN
满分1
参评模型6

模型排名

名次 模型 机构 分数 来源
1 Gemini 2.5 Pro Google 93.0 来源 ↗
2 Gemini 1.5 Pro Google 82.6 来源 ↗
3 Gemini 1.5 Flash Google 71.9 来源 ↗
4 Gemini 2.0 Flash Google 69.2 来源 ↗
5 Gemini 1.5 Flash 8B Google 54.7 来源 ↗
6 Gemini 2.5 Flash Google 32.0 来源 ↗