Benchmark

OJBench

reasoning text

OJBench is a competition-level code benchmark designed to assess the competitive-level code reasoning abilities of large language models. It comprises 232 programming competition problems from NOI and ICPC, categorized into Easy, Medium, and Hard difficulty levels. The benchmark evaluates models' ability to solve complex competitive programming challenges using Python and C++.

语言EN
满分1
参评模型4

模型排名

名次 模型 机构 分数 来源
1 Qwen3-235B-A22B-Thinking-2507 Alibaba Cloud / Qwen Team 32.5 来源 ↗
2 Qwen3-Next-80B-A3B-Thinking Alibaba Cloud / Qwen Team 29.7 来源 ↗
3 Kimi K2 Instruct Moonshot AI 27.1 来源 ↗
4 Kimi K2-Instruct-0905 Moonshot AI 27.1 来源 ↗