Benchmark
HumanEval+
reasoning
text
Enhanced version of HumanEval that extends the original test cases by 80x using EvalPlus framework for rigorous evaluation of LLM-synthesized code functional correctness, detecting previously undetected wrong code
语言EN
满分1
参评模型8
模型排名
| 名次 | 模型 | 机构 | 分数 | 来源 |
|---|---|---|---|---|
| 1 | Phi 4 Reasoning | Microsoft | 92.9 | 来源 ↗ |
| 2 | Phi 4 Reasoning Plus | Microsoft | 92.3 | 来源 ↗ |
| 3 | Granite 3.3 8B Base | IBM | 86.1 | 来源 ↗ |
| 4 | Granite 3.3 8B Instruct | IBM | 86.1 | 来源 ↗ |
| 5 | Phi 4 | Microsoft | 82.8 | 来源 ↗ |
| 6 | IBM Granite 4.0 Tiny Preview | IBM | 78.3 | 来源 ↗ |
| 7 | Qwen2.5 32B Instruct | Alibaba Cloud / Qwen Team | 52.4 | 来源 ↗ |
| 8 | Qwen2.5 14B Instruct | Alibaba Cloud / Qwen Team | 51.2 | 来源 ↗ |