Benchmark

MBPP EvalPlus

reasoning general text

MBPP (Mostly Basic Python Problems) is a benchmark of 974 crowd-sourced Python programming problems designed to be solvable by entry-level programmers. EvalPlus extends MBPP with significantly more test cases (35x) for more rigorous evaluation of LLM-synthesized code, providing high-quality and precise evaluation.

语言EN
满分1
参评模型2

模型排名

名次 模型 机构 分数 来源
1 Llama 3.1 405B Instruct Meta 88.6 来源 ↗
2 Llama 3.3 70B Instruct Meta 87.6 来源 ↗