modelbenchmark.io
All benchmarks

Aider polyglot

Coding · fixed question set

225 Exercism exercises across C++, Go, Java, JavaScript, Python and Rust, edited in place through the aider tool.

How far to trust it

⚠️ Not maintained past 2025-10-03, so no recent model appears. It measures editing format compliance as much as coding: a model that writes correct code in the wrong diff format scores zero.

Measured

Models scored
60
Spread between models
23.6
standard deviation, points
Measurement noise
2.7
inferred from other benchmarks
Weight
0.99
share of spread that is signal

The weight is (spread² − noise²) / spread². It is computed, never chosen. Most of this benchmark's spread is real, so it still separates models.

Every model scored

the median across configurations, so one heroic run cannot lead

View
RankModelLabScore
1GPT 5OpenAI86.7
2o3-pro84.9
3gemini-2.5-pro-preview-06-0583.1
4Grok 4xAI79.6
5gemini-2.5-pro-preview-06-0579.1
6o3OpenAI79.1
7o3 (high) + gpt-4.178.2
8Gemini 2.5 Pro Preview 05-0676.9
9Gemini 2.5 Pro Preview 03-2572.9
10DeepSeek-V3.2-ExpDeepSeek72.2
11Claude Opus 4Anthropic72.0
12o4-miniOpenAI72.0
13DeepSeek R171.4
14claude-opus-4-2025051470.7
15Claude 3.7 SonnetAnthropic64.9
16DeepSeek R1 + claude-3-5-sonnet-2024102264.0
17o1-2024-12-17OpenAI61.7
18Claude Sonnet 4Anthropic61.3
19claude-3-7-sonnet-2025021960.4
20Qwen3 235B A22B diff, no think, Alibaba API59.6
21Kimi K2 InstructKimi59.1
22O3 MiniOpenAI57.1
23DeepSeek R1DeepSeek56.9
24claude-sonnet-4-2025051456.4
25DeepSeek V355.1
26gemini-2.5-flash-preview-05-2055.1
27Quasar Alpha54.7
28Grok 3 BetaxAI53.3
29Optimus Alpha52.9
30GPT 4.1OpenAI52.4
31Claude 3.5 SonnetAnthropic51.6
32DeepSeek Chat V348.4
33gemini-2.5-flash-preview-04-1747.1
34GPT-4.5OpenAI44.9
35gemini-2.5-flash-preview-05-2044.0
36Grok 3 Mini Beta42.0
37gpt-oss-120bOpenAI41.8
38Qwen3 32BAlibaba40.0
39gemini-exp-120638.2
40chatgpt-4o-latest36.2
41Gemini 2.0 Pro exp-02-0535.6
42o1-miniOpenAI32.9
43GPT 4.1 miniOpenAI32.4
44Claude 3.5 HaikuAnthropic28.0
45QwQ-32B + Qwen 2.5 Coder Instruct26.2
46Gemini 2.0 Flash ThinkingGoogle22.2
47qwen-max-2025-01-2521.8
48QwQ 32BAlibaba20.9
49GPT 4oOpenAI20.6
50gemini-2.0-flash-thinking-exp-01-2118.2
51DeepSeek Chat V2.517.8
52Llama 4 Maverick InstructMeta15.6
53yi-lightning12.9
54Qwen2.5 Coder 32B InstructAlibaba12.2
55command-a-03-2025-quality12.0
56Codestral 25.0111.1
57openhands-lm-32b-v0.110.2
58GPT-4.1 nanoOpenAI8.9
59Gemma 3 27BGoogle4.9
60GPT-4o miniOpenAI3.6

The benchmark's own page