Benchmark

HumanEval+

reasoning text

Enhanced version of HumanEval that extends the original test cases by 80x using EvalPlus framework for rigorous evaluation of LLM-synthesized code functional correctness, detecting previously undetected wrong code

语言EN
满分1
参评模型8

模型排名

名次 模型 机构 分数 来源
1 Phi 4 Reasoning Microsoft 92.9 来源 ↗
2 Phi 4 Reasoning Plus Microsoft 92.3 来源 ↗
3 Granite 3.3 8B Base IBM 86.1 来源 ↗
4 Granite 3.3 8B Instruct IBM 86.1 来源 ↗
5 Phi 4 Microsoft 82.8 来源 ↗
6 IBM Granite 4.0 Tiny Preview IBM 78.3 来源 ↗
7 Qwen2.5 32B Instruct Alibaba Cloud / Qwen Team 52.4 来源 ↗
8 Qwen2.5 14B Instruct Alibaba Cloud / Qwen Team 51.2 来源 ↗