Benchmark

MMLU-redux-2.0

language reasoning math general text

A curated version of the MMLU benchmark featuring manually re-annotated 5,700 questions across 57 subjects to identify and correct errors in the original dataset. Addresses the 6.49% error rate found in MMLU and provides more reliable evaluation metrics for language models.

语言EN
满分1
参评模型1

模型排名

名次 模型 机构 分数 来源
1 Kimi K2 Base Moonshot AI 90.2 来源 ↗