Benchmark

Aider-Polyglot Edit

general code text

A challenging multi-language coding benchmark that evaluates models' code editing abilities across C++, Go, Java, JavaScript, Python, and Rust. Contains 225 of Exercism's most difficult programming problems, selected as problems that were solved by 3 or fewer out of 7 top coding models. The benchmark focuses on code editing tasks and measures both correctness of solutions and proper edit format usage. Designed to re-calibrate evaluation scales so top models score between 5-50%.

语言EN
满分1
参评模型10

模型排名

名次 模型 机构 分数 来源
1 DeepSeek-V3 DeepSeek 79.7 来源 ↗
2 Gemini 2.5 Pro Google 72.7 来源 ↗
3 o3-mini OpenAI 60.4 来源 ↗
4 o4-mini OpenAI 58.2 来源 ↗
5 Gemini 2.5 Flash Google 56.7 来源 ↗
6 GPT-4.1 OpenAI 52.9 来源 ↗
7 GPT-4.5 OpenAI 44.9 来源 ↗
8 GPT-4.1 mini OpenAI 31.6 来源 ↗
9 GPT-4o OpenAI 18.2 来源 ↗
10 GPT-4.1 nano OpenAI 6.2 来源 ↗