Benchmark

Multipl-E MBPP

general reasoning text 多语言

MultiPL-E extends the Mostly Basic Python Problems (MBPP) benchmark to 18+ programming languages for evaluating multilingual code generation capabilities. MBPP contains 974 crowd-sourced programming problems designed to be solvable by entry-level programmers, covering programming fundamentals and standard library functionality. Each problem includes a task description, code solution, and automated test cases.

语言EN
满分1
参评模型3

模型排名

名次 模型 机构 分数 来源
1 Llama 3.1 405B Instruct Meta 65.7 来源 ↗
2 Llama 3.1 70B Instruct Meta 62.0 来源 ↗
3 Llama 3.1 8B Instruct Meta 52.4 来源 ↗