Benchmark
HumanEval-Average
reasoning
text
A variant of the HumanEval benchmark that measures functional correctness for synthesizing programs from docstrings, consisting of 164 original programming problems assessing language comprehension, algorithms, and simple mathematics
语言EN
满分1
参评模型1
模型排名
| 名次 | 模型 | 机构 | 分数 | 来源 |
|---|---|---|---|---|
| 1 | Codestral-22B | Mistral AI | 61.5 | 来源 ↗ |