Benchmark
SWE-bench Verified (Agentic Coding)
reasoning
code
text
SWE-bench Verified is a human-filtered subset of 500 software engineering problems drawn from real GitHub issues across 12 popular Python repositories. Given a codebase and an issue description, language models are tasked with generating patches that resolve the described problems. This benchmark evaluates AI's real-world agentic coding skills by requiring models to navigate complex codebases, understand software engineering problems, and coordinate changes across multiple functions, classes, and files to fix well-defined issues with clear descriptions.
语言EN
满分1
参评模型2
模型排名
| 名次 | 模型 | 机构 | 分数 | 来源 |
|---|---|---|---|---|
| 1 | Claude Sonnet 4.5 | Anthropic | 77.2 | 来源 ↗ |
| 2 | Kimi K2 Instruct | Moonshot AI | 65.8 | 来源 ↗ |