Benchmark
OpenAI-MRCR: 2 needle 128k
long_context
reasoning
text
Multi-round Co-reference Resolution (MRCR) benchmark for evaluating an LLM's ability to distinguish between multiple needles hidden in long context. Models are given a long, multi-turn synthetic conversation and must retrieve a specific instance of a repeated request, requiring reasoning and disambiguation skills beyond simple retrieval.
语言EN
满分1
参评模型7