Benchmark

HumanEval-Mul

reasoning text 多语言

A multilingual variant of the HumanEval benchmark that measures functional correctness for synthesizing programs from docstrings, consisting of 164 original programming problems assessing language comprehension, algorithms, and simple mathematics

语言EN
满分1
参评模型2

模型排名

名次 模型 机构 分数 来源
1 DeepSeek-V3 DeepSeek 82.6 来源 ↗
2 DeepSeek-V2.5 DeepSeek 73.8 来源 ↗