Benchmark

CruxEval-O

reasoning text

CruxEval-O is the output prediction task of the CRUXEval benchmark, designed to evaluate code reasoning, understanding, and execution capabilities. It consists of 800 Python functions (3-13 lines) where models must predict the output given a function and input. The benchmark tests fundamental code execution reasoning abilities and goes beyond simple code generation to assess deeper understanding of program behavior.

语言EN
满分1
参评模型1

模型排名

名次 模型 机构 分数 来源
1 Codestral-22B Mistral AI 51.3 来源 ↗