Benchmark
Aider-Polyglot Edit
general
code
text
A challenging multi-language coding benchmark that evaluates models' code editing abilities across C++, Go, Java, JavaScript, Python, and Rust. Contains 225 of Exercism's most difficult programming problems, selected as problems that were solved by 3 or fewer out of 7 top coding models. The benchmark focuses on code editing tasks and measures both correctness of solutions and proper edit format usage. Designed to re-calibrate evaluation scales so top models score between 5-50%.
语言EN
满分1
参评模型10
模型排名
| 名次 | 模型 | 机构 | 分数 | 来源 |
|---|---|---|---|---|
| 1 | DeepSeek-V3 | DeepSeek | 79.7 | 来源 ↗ |
| 2 | Gemini 2.5 Pro | 72.7 | 来源 ↗ | |
| 3 | o3-mini | OpenAI | 60.4 | 来源 ↗ |
| 4 | o4-mini | OpenAI | 58.2 | 来源 ↗ |
| 5 | Gemini 2.5 Flash | 56.7 | 来源 ↗ | |
| 6 | GPT-4.1 | OpenAI | 52.9 | 来源 ↗ |
| 7 | GPT-4.5 | OpenAI | 44.9 | 来源 ↗ |
| 8 | GPT-4.1 mini | OpenAI | 31.6 | 来源 ↗ |
| 9 | GPT-4o | OpenAI | 18.2 | 来源 ↗ |
| 10 | GPT-4.1 nano | OpenAI | 6.2 | 来源 ↗ |