Benchmark
OJBench
reasoning
text
OJBench is a competition-level code benchmark designed to assess the competitive-level code reasoning abilities of large language models. It comprises 232 programming competition problems from NOI and ICPC, categorized into Easy, Medium, and Hard difficulty levels. The benchmark evaluates models' ability to solve complex competitive programming challenges using Python and C++.
语言EN
满分1
参评模型4
模型排名
| 名次 | 模型 | 机构 | 分数 | 来源 |
|---|---|---|---|---|
| 1 | Qwen3-235B-A22B-Thinking-2507 | Alibaba Cloud / Qwen Team | 32.5 | 来源 ↗ |
| 2 | Qwen3-Next-80B-A3B-Thinking | Alibaba Cloud / Qwen Team | 29.7 | 来源 ↗ |
| 3 | Kimi K2 Instruct | Moonshot AI | 27.1 | 来源 ↗ |
| 4 | Kimi K2-Instruct-0905 | Moonshot AI | 27.1 | 来源 ↗ |