Benchmark

MRCR 1M

long_context reasoning general text

MRCR 1M is a variant of the Multi-Round Coreference Resolution benchmark designed for testing extremely long context capabilities with approximately 1 million tokens. It evaluates models' ability to maintain reasoning and attention across ultra-long conversations.

语言EN
满分1
参评模型1

模型排名

名次 模型 机构 分数 来源
1 Gemini 2.0 Flash-Lite Google 58.0 来源 ↗