modelbenchmark.io
All benchmarks

SWE-bench Verified

Coding · fixed question set

500 real GitHub issues from Python repositories, human-checked to be solvable. The model must produce a patch that passes the repository's own tests.

How far to trust it

🚨 A SWE-bench number is never a fact about a model alone. It is a fact about a model inside an agent scaffold, and the scaffold changes the score. Every figure here records which scaffold ran. One audit of the top entries found about 19.78% of 'solved' cases were semantically wrong: the patch passed the tests without fixing the issue. Read this benchmark as a floor on capability, never as a percentage of real bugs fixed.

Measured

Models scored
84
Spread between models
20.4
standard deviation, points
Measurement noise
2.0
published stderr
Weight
0.99
share of spread that is signal

The weight is (spread² − noise²) / spread². It is computed, never chosen. Most of this benchmark's spread is real, so it still separates models.

Every model scored

the median across configurations, so one heroic run cannot lead

View
RankModelLabScore
1Claude Opus 4.7Anthropic83.5
2GPT-5.5OpenAI80.6
3Doubao-Seed-CodeByteDance78.8
4GLM-5.2Zhipu78.7
5DeepSeek v4DeepSeek77.6
6Qwen3.7-MaxAlibaba77.3
7Claude Opus 4.6Anthropic77.2
8GPT-5.4OpenAI76.9
9Claude 4.5 OpusAnthropic76.8
10Claude Opus 4.5Anthropic76.7
11Kimi K2.6Kimi76.7
12Qwen3.6 Max PreviewAlibaba76.7
13MiniMax M2.5MiniMax75.8
14Gemini 3.1 Pro PreviewGoogle75.6
15Gemini 3 Flash PreviewGoogle75.6
16Claude 4.6 OpusAnthropic75.6
17Gemini 3.5 FlashGoogle75.6
18Claude Sonnet 4.6Anthropic75.2
19GPT-5.3 CodexOpenAI74.8
20GLM-5.1Zhipu74.2
21Claude Opus 4.1Anthropic73.3
22Gemini 3 Pro PreviewGoogle72.9
23GPT 5.2 CodexOpenAI72.8
24GLM 5Zhipu72.4
25GPT 5.2OpenAI72.3
26Kimi K2.5Kimi72.3
27Claude 4.5 SonnetAnthropic72.0
28GPT 5OpenAI71.6
29Claude Sonnet 4.5Anthropic71.3
30Claude 4 SonnetAnthropic71.0
31Claude Opus 4Anthropic70.7
32Claude 4 OpusAnthropic70.4
33Claude 4.5 HaikuAnthropic66.6
34GPT 5.1OpenAI66.5
35GPT 5.1 CodexOpenAI66.0
36Claude Sonnet 4Anthropic65.9
37DeepSeek-V3.2-ExpDeepSeek65.0
38o1-2024-12-17OpenAI64.6
39MultipleAnthropic64.1
40Claude 3.7 Sonnet w/ Review HeavyAnthropic62.4
41swe-searchAnthropic62.2
42MiniMax M2MiniMax61.0
43o3OpenAI60.4
44Claude 3.7 SonnetAnthropic60.3
45GPT 5 miniOpenAI59.8
46Kimi K2 InstructKimi59.4
47TTSOpenAI58.8
48Qwen3.6 PlusAlibaba57.9
49Devstral SmallMistral56.4
50Qwen3-Coder-30B-A3B-InstructAlibaba56.3
51Gemini 2.5 ProGoogle55.6
52GLM 4.6Zhipu55.4
53Qwen3-Coder 480B/A35B InstructAlibaba55.4
54o4-miniOpenAI54.8
55GLM 4.5Zhipu54.2
56DevstralMistral53.8
57Frogboss 32B 2510Microsoft53.6
58Claude 3.5 SonnetAnthropic50.7
59TTSAlibaba47.0
60DevStral Small 2505Mistral46.8
61Frogmini 14B 2510Microsoft45.0
62GPT 4.1OpenAI44.1
63Amazon.nova Premier v1:0Amazon42.4
64O3 MiniOpenAI42.4
65DeepSWE-PreviewAgentica42.2
66Llama3-SWE-RL-70BMeta41.2
67Claude 3.5 HaikuAnthropic40.6
68SWE-agent-LM-32BAlibaba40.2
69DevStral Small 2507Mistral38.0
70GPT 5 nanoOpenAI34.8
71Qwen2.5Alibaba31.5
72GPT 4oOpenAI30.0
73Gemini 2.0 Flash ThinkingGoogle28.9
74Gemini 2.5 FlashGoogle28.7
75gpt-oss-120bOpenAI26.0
76GPT 4.1 miniOpenAI23.9
77Qwen2.5 Coder 32B InstructAlibaba23.5
78MCTS Refine 7B23.2
79Llama 4 Maverick InstructMeta21.0
80GPT-4OpenAI12.6
81Claude 3 OpusAnthropic11.4
82Llama 4 Scout InstructMeta9.1
83Claude 2Anthropic4.4
84SWE-Llama 7BMeta1.4

The benchmark's own page