Benchmark
HumanEval-Mul
reasoning
text
多语言
A multilingual variant of the HumanEval benchmark that measures functional correctness for synthesizing programs from docstrings, consisting of 164 original programming problems assessing language comprehension, algorithms, and simple mathematics
语言EN
满分1
参评模型2
模型排名
| 名次 | 模型 | 机构 | 分数 | 来源 |
|---|---|---|---|---|
| 1 | DeepSeek-V3 | DeepSeek | 82.6 | 来源 ↗ |
| 2 | DeepSeek-V2.5 | DeepSeek | 73.8 | 来源 ↗ |