Benchmark

MRCR v2 (8-needle)

long_context reasoning general text

MRCR v2 (8-needle) is a variant of the Multi-Round Coreference Resolution benchmark that includes 8 needle items to retrieve from long contexts. This tests models' ability to simultaneously track and reason about multiple pieces of information across extended conversations.

语言EN
满分1
参评模型1

模型排名

名次 模型 机构 分数 来源
1 Gemini 2.5 Pro Preview 06-05 Google 16.4 来源 ↗