Benchmark

OpenAI-MRCR: 2 needle 1M

long_context reasoning text

Multi-Round Co-reference Resolution benchmark that tests an LLM's ability to distinguish between multiple similar needles hidden in long conversations. Models must reproduce specific instances of content (e.g., 'Return the 2nd poem about tapirs') from multi-turn synthetic conversations, requiring reasoning about context, ordering, and subtle differences between similar outputs.

语言EN
满分1
参评模型3

模型排名

名次 模型 机构 分数 来源
1 GPT-4.1 OpenAI 46.3 来源 ↗
2 GPT-4.1 mini OpenAI 33.3 来源 ↗
3 GPT-4.1 nano OpenAI 12.0 来源 ↗