Benchmark

SWE-bench Verified (Agentless)

general reasoning text

A human-validated subset of SWE-bench that evaluates language models' ability to resolve real-world GitHub issues using an agentless approach. The benchmark tests models on software engineering problems requiring understanding and coordinating changes across multiple functions, classes, and files simultaneously.

语言EN
满分1
参评模型1

模型排名

名次 模型 机构 分数 来源
1 Kimi K2 Instruct Moonshot AI 51.8 来源 ↗