Aider polyglot
Coding · fixed question set
225 Exercism exercises across C++, Go, Java, JavaScript, Python and Rust, edited in place through the aider tool.
How far to trust it
⚠️ Not maintained past 2025-10-03, so no recent model appears. It measures editing format compliance as much as coding: a model that writes correct code in the wrong diff format scores zero.
Measured
Models scored
60
Spread between models
23.6
standard deviation, points
Measurement noise
2.7
inferred from other benchmarks
Weight
0.99
share of spread that is signal
The weight is (spread² − noise²) / spread². It is computed, never chosen. Most of this benchmark's spread is real, so it still separates models.
Every model scored
the median across configurations, so one heroic run cannot lead
| Rank | Model | Lab | Score |
|---|---|---|---|
| 1 | GPT 5 | OpenAI | 86.7 |
| 2 | o3-pro | — | 84.9 |
| 3 | gemini-2.5-pro-preview-06-05 | — | 83.1 |
| 4 | Grok 4 | xAI | 79.6 |
| 5 | gemini-2.5-pro-preview-06-05 | — | 79.1 |
| 6 | o3 | OpenAI | 79.1 |
| 7 | o3 (high) + gpt-4.1 | — | 78.2 |
| 8 | Gemini 2.5 Pro Preview 05-06 | — | 76.9 |
| 9 | Gemini 2.5 Pro Preview 03-25 | — | 72.9 |
| 10 | DeepSeek-V3.2-Exp | DeepSeek | 72.2 |
| 11 | Claude Opus 4 | Anthropic | 72.0 |
| 12 | o4-mini | OpenAI | 72.0 |
| 13 | DeepSeek R1 | — | 71.4 |
| 14 | claude-opus-4-20250514 | — | 70.7 |
| 15 | Claude 3.7 Sonnet | Anthropic | 64.9 |
| 16 | DeepSeek R1 + claude-3-5-sonnet-20241022 | — | 64.0 |
| 17 | o1-2024-12-17 | OpenAI | 61.7 |
| 18 | Claude Sonnet 4 | Anthropic | 61.3 |
| 19 | claude-3-7-sonnet-20250219 | — | 60.4 |
| 20 | Qwen3 235B A22B diff, no think, Alibaba API | — | 59.6 |
| 21 | Kimi K2 Instruct | Kimi | 59.1 |
| 22 | O3 Mini | OpenAI | 57.1 |
| 23 | DeepSeek R1 | DeepSeek | 56.9 |
| 24 | claude-sonnet-4-20250514 | — | 56.4 |
| 25 | DeepSeek V3 | — | 55.1 |
| 26 | gemini-2.5-flash-preview-05-20 | — | 55.1 |
| 27 | Quasar Alpha | — | 54.7 |
| 28 | Grok 3 Beta | xAI | 53.3 |
| 29 | Optimus Alpha | — | 52.9 |
| 30 | GPT 4.1 | OpenAI | 52.4 |
| 31 | Claude 3.5 Sonnet | Anthropic | 51.6 |
| 32 | DeepSeek Chat V3 | — | 48.4 |
| 33 | gemini-2.5-flash-preview-04-17 | — | 47.1 |
| 34 | GPT-4.5 | OpenAI | 44.9 |
| 35 | gemini-2.5-flash-preview-05-20 | — | 44.0 |
| 36 | Grok 3 Mini Beta | — | 42.0 |
| 37 | gpt-oss-120b | OpenAI | 41.8 |
| 38 | Qwen3 32B | Alibaba | 40.0 |
| 39 | gemini-exp-1206 | — | 38.2 |
| 40 | chatgpt-4o-latest | — | 36.2 |
| 41 | Gemini 2.0 Pro exp-02-05 | — | 35.6 |
| 42 | o1-mini | OpenAI | 32.9 |
| 43 | GPT 4.1 mini | OpenAI | 32.4 |
| 44 | Claude 3.5 Haiku | Anthropic | 28.0 |
| 45 | QwQ-32B + Qwen 2.5 Coder Instruct | — | 26.2 |
| 46 | Gemini 2.0 Flash Thinking | 22.2 | |
| 47 | qwen-max-2025-01-25 | — | 21.8 |
| 48 | QwQ 32B | Alibaba | 20.9 |
| 49 | GPT 4o | OpenAI | 20.6 |
| 50 | gemini-2.0-flash-thinking-exp-01-21 | — | 18.2 |
| 51 | DeepSeek Chat V2.5 | — | 17.8 |
| 52 | Llama 4 Maverick Instruct | Meta | 15.6 |
| 53 | yi-lightning | — | 12.9 |
| 54 | Qwen2.5 Coder 32B Instruct | Alibaba | 12.2 |
| 55 | command-a-03-2025-quality | — | 12.0 |
| 56 | Codestral 25.01 | — | 11.1 |
| 57 | openhands-lm-32b-v0.1 | — | 10.2 |
| 58 | GPT-4.1 nano | OpenAI | 8.9 |
| 59 | Gemma 3 27B | 4.9 | |
| 60 | GPT-4o mini | OpenAI | 3.6 |