Benchmark

MBPP EvalPlus (base)

reasoning general text

MBPP (Mostly Basic Python Problems) is a benchmark of 974 crowd-sourced Python programming problems designed to be solvable by entry-level programmers. EvalPlus extends MBPP with significantly more test cases (35x) for more rigorous evaluation of LLM-synthesized code, providing high-quality and precise evaluation.

语言EN
满分1
参评模型1

模型排名

名次 模型 机构 分数 来源
1 Llama 3.1 8B Instruct Meta 72.8 来源 ↗