Benchmark
Aider-Polyglot
general
code
text
A coding benchmark that evaluates LLMs on 225 challenging Exercism programming exercises across C++, Go, Java, JavaScript, Python, and Rust. Models receive two attempts to solve each problem, with test error feedback provided after the first attempt if it fails. The benchmark measures both initial problem-solving ability and capacity to edit code based on error feedback, providing an end-to-end evaluation of code generation and editing capabilities across multiple programming languages.
语言EN
满分1
参评模型31
模型排名
| 名次 | 模型 | 机构 | 分数 | 来源 |
|---|---|---|---|---|
| 1 | GPT-5 | OpenAI | 88.0 | 来源 ↗ |
| 2 | Gemini 2.5 Pro Preview 06-05 | 82.2 | 来源 ↗ | |
| 3 | o3 | OpenAI | 81.3 | 来源 ↗ |
| 4 | Gemini 2.5 Pro | 76.5 | 来源 ↗ | |
| 5 | DeepSeek-V3.2-Exp | DeepSeek | 74.5 | 来源 ↗ |
| 6 | DeepSeek Reasoner | DeepSeek | 74.2 | 来源 ↗ |
| 7 | Claude Opus 4 (latest) | Anthropic | 72.0 | 来源 ↗ |
| 8 | DeepSeek-R1-0528 | DeepSeek | 71.6 | 来源 ↗ |
| 9 | DeepSeek Chat | DeepSeek | 70.2 | 来源 ↗ |
| 10 | o4-mini | OpenAI | 68.9 | 来源 ↗ |
| 11 | DeepSeek-V3.1 | DeepSeek | 68.4 | 来源 ↗ |
| 12 | o3-mini | OpenAI | 66.7 | 来源 ↗ |
| 13 | Gemini 2.5 Flash | 61.9 | 来源 ↗ | |
| 14 | Claude Sonnet 4 (latest) | Anthropic | 61.3 | 来源 ↗ |
| 15 | Kimi K2 Instruct | Moonshot AI | 60.0 | 来源 ↗ |
| 16 | Kimi K2-Instruct-0905 | Moonshot AI | 60.0 | 来源 ↗ |
| 17 | Qwen3-235B-A22B-Instruct-2507 | Alibaba Cloud / Qwen Team | 57.3 | 来源 ↗ |
| 18 | GPT-4.1 | OpenAI | 51.6 | 来源 ↗ |
| 19 | Qwen3-Next-80B-A3B-Instruct | Alibaba Cloud / Qwen Team | 49.8 | 来源 ↗ |
| 20 | DeepSeek-V3 | DeepSeek | 49.6 | 来源 ↗ |
| 21 | Magistral Medium | Mistral AI | 47.1 | 来源 ↗ |
| 22 | GPT-4.1 mini | OpenAI | 34.7 | 来源 ↗ |
| 23 | GPT-4o | OpenAI | 30.7 | 来源 ↗ |
| 24 | Gemini 2.5 Flash-Lite | 26.7 | 来源 ↗ | |
| 25 | GPT-4o | OpenAI | 23.1 | 来源 ↗ |
| 26 | Qwen Max | Alibaba Cloud / Qwen Team | 21.8 | 来源 ↗ |
| 27 | GPT-4o (2024-11-20) | OpenAI | 18.2 | 来源 ↗ |
| 28 | Llama 4 Maverick 17B Instruct | Meta | 15.6 | 来源 ↗ |
| 29 | Command A | Cohere | 12.0 | 来源 ↗ |
| 30 | Codestral (latest) | Mistral AI | 11.1 | 来源 ↗ |
| 31 | GPT-4.1 nano | OpenAI | 9.8 | 来源 ↗ |