Benchmark

HumanEvalFIM-Average

general text

Average evaluation of HumanEval Fill-in-the-Middle benchmark variants (single-line, multi-line, random-span) for assessing code infilling capabilities of language models

语言EN
满分1
参评模型1

模型排名

名次 模型 机构 分数 来源
1 Codestral-22B Mistral AI 91.6 来源 ↗