Benchmark
MRCR
long_context
reasoning
general
text
MRCR (Multi-Round Coreference Resolution) is a synthetic long-context reasoning task where models must navigate long conversations to reproduce specific model outputs. It tests the ability to distinguish between similar requests and reason about ordering while maintaining attention across extended contexts.
语言EN
满分1
参评模型6
模型排名
| 名次 | 模型 | 机构 | 分数 | 来源 |
|---|---|---|---|---|
| 1 | Gemini 2.5 Pro | 93.0 | 来源 ↗ | |
| 2 | Gemini 1.5 Pro | 82.6 | 来源 ↗ | |
| 3 | Gemini 1.5 Flash | 71.9 | 来源 ↗ | |
| 4 | Gemini 2.0 Flash | 69.2 | 来源 ↗ | |
| 5 | Gemini 1.5 Flash 8B | 54.7 | 来源 ↗ | |
| 6 | Gemini 2.5 Flash | 32.0 | 来源 ↗ |