Benchmark

OpenAI-MRCR: 2 needle 128k

long_context reasoning text

Multi-round Co-reference Resolution (MRCR) benchmark for evaluating an LLM's ability to distinguish between multiple needles hidden in long context. Models are given a long, multi-turn synthetic conversation and must retrieve a specific instance of a repeated request, requiring reasoning and disambiguation skills beyond simple retrieval.

语言EN
满分1
参评模型7

模型排名

名次 模型 机构 分数 来源
1 GPT-5 OpenAI 95.2 来源 ↗
2 GPT-4.1 OpenAI 57.2 来源 ↗
3 GPT-4.1 mini OpenAI 47.2 来源 ↗
4 GPT-4.5 OpenAI 38.5 来源 ↗
5 GPT-4.1 nano OpenAI 36.6 来源 ↗
6 GPT-4o OpenAI 31.9 来源 ↗
7 o3-mini OpenAI 18.7 来源 ↗