Benchmark

OmniMath

math reasoning text

A Universal Olympiad Level Mathematic Benchmark for Large Language Models containing 4,428 competition-level problems with rigorous human annotation, categorized into over 33 sub-domains and spanning more than 10 distinct difficulty levels

语言EN
满分1
参评模型2

模型排名

名次 模型 机构 分数 来源
1 Phi 4 Reasoning Plus Microsoft 81.9 来源 ↗
2 Phi 4 Reasoning Microsoft 76.6 来源 ↗