Benchmark

RULER

long_context reasoning text

RULER (What's the Real Context Size of Your Long-Context Language Models?) is a synthetic benchmark designed to comprehensively evaluate the long-context capabilities of language models. It expands on needle-in-a-haystack (NIAH) testing by introducing new task categories including multi-hop tracing and aggregation tasks. The benchmark provides flexible configurations for customized sequence length and task complexity, evaluating 17 long-context language models across 13 representative tasks to reveal that despite models claiming 32K+ token context sizes, only half maintain satisfactory performance at 32K length.

语言EN
满分1
参评模型2

模型排名

名次 模型 机构 分数 来源
1 Phi-3.5-MoE-instruct Microsoft 87.1 来源 ↗
2 Phi-3.5-mini-instruct Microsoft 84.1 来源 ↗