Benchmark

Multipl-E HumanEval

language general text 多语言

MultiPL-E is a scalable and extensible approach to benchmarking neural code generation that translates unit test-driven code generation benchmarks across multiple programming languages. It extends the HumanEval benchmark to 18 additional programming languages, enabling evaluation of code generation models across diverse programming paradigms and providing insights into how models generalize programming knowledge across language boundaries.

语言EN
满分1
参评模型3

模型排名

名次 模型 机构 分数 来源
1 Llama 3.1 405B Instruct Meta 75.2 来源 ↗
2 Llama 3.1 70B Instruct Meta 65.5 来源 ↗
3 Llama 3.1 8B Instruct Meta 50.8 来源 ↗