modelbenchmark.io
All benchmarks

GPQA Diamond

Reasoning · fixed question set

448 graduate-level questions in biology, physics and chemistry, written so that a skilled non-expert with web access still scores near 34%.

How far to trust it

The most reported benchmark on this site, and the best understood. Epoch AI measured the gap between what labs report and what an independent run produces at about +0.7 points on average, inside their confidence intervals. On this benchmark a self-reported number is worth roughly what an independent one is. The set is fixed and public, so contamination rises over time.

Measured

Models scored
219
Spread between models
22.8
standard deviation, points
Measurement noise
2.5
published stderr
Weight
0.99
share of spread that is signal

The weight is (spread² − noise²) / spread². It is computed, never chosen. Most of this benchmark's spread is real, so it still separates models.

Every model scored

the median across configurations, so one heroic run cannot lead

View
RankModelLabScore
1GPT-6 AstraOpenAI95.8
2Gemini 3.8 FlashGoogle95.4
3Gemini 3.7 FlashGoogle94.8
4GPT-5.4 ProOpenAI94.6
5Gemini 3.1 Pro PreviewGoogle94.3
6GPT-5.5 Pro pre-releaseOpenAI93.9
7Grok 4.6xAI93.6
8Grok 4.5xAI93.4
9Claude Opus 5Anthropic92.9
10Qwen3.8 MaxAlibaba92.7
11Gemini 3 Pro PreviewGoogle92.6
12GPT-5.5OpenAI92.3
13Qwen3.8 Max (0902)Alibaba92.3
14Kimi K3Kimi91.9
15GPT-5.6 SolOpenAI91.7
16DeepSeek V4 Pro 0813DeepSeek91.7
17DeepSeek V4 Flash 0731DeepSeek91.0
18GLM-5.3Zhipu90.9
19MiniMax M3MiniMax90.9
20Qwen3.7-MaxAlibaba90.9
21Kimi K2.6Kimi90.8
22GPT-5.6 TerraOpenAI90.3
23DeepSeek v4DeepSeek90.3
24GLM-5.3-FlashZhipu90.2
25GLM-5.1Zhipu89.9
26GLM-5.2Zhipu89.9
27Muse SparkMeta89.8
28Claude Opus 4.8Anthropic89.7
29GPT-5.4OpenAI89.4
30Grok 4.20xAI89.3
31Gemini 3.5 FlashGoogle88.9
32Grok 4.3 BetaxAI88.8
33Claude Opus 4.6Anthropic88.8
34Inkling SmallThinking Machines88.5
35Qwen3.6 PlusAlibaba88.4
36InklingThinking Machines88.3
37Claude Opus 4.7Anthropic88.3
38GPT 5.2OpenAI88.0
39Kimi K2.7 CodeKimi87.9
40Qwen3.7 PlusAlibaba87.9
41GLM 5Zhipu87.8
42Kimi K2.5Kimi87.6
43Qwen3.6 Max PreviewAlibaba87.4
44Grok 4xAI87.0
45GPT-5.6 LunaOpenAI87.0
46Gemini 3.6 FlashGoogle86.4
47Qwen3.5 397B-A17BAlibaba86.4
48GPT 5.1OpenAI86.3
49Gemini 3 Flash PreviewGoogle86.3
50Qwen3.5 397B-A17BAlibaba85.9
51Qwen3.6 27BAlibaba85.9
52Claude Opus 4.5Anthropic85.8
53Claude Sonnet 5Anthropic85.4
54Claude Opus 4.8Anthropic85.3
55GPT 5OpenAI85.3
56Nemotron 3 UltraNVIDIA85.3
57Gemini 2.5 ProGoogle85.3
58GPT-5.4 miniOpenAI85.2
59Qwen3.5 PlusAlibaba84.8
60Qwen3.6 27BAlibaba84.8
61Qwen3.6 35B-A3BAlibaba84.8
62Kimi K2 InstructKimi84.2
63Gemini 2.5 Pro ExpGoogle83.8
64Qwen3.6 35B-A3BAlibaba83.8
65Claude Fable 5Anthropic83.3
66Claude Sonnet 4.6Anthropic83.3
67GLM-4.7Zhipu83.3
68Qwen3.6 FlashAlibaba83.3
69GPT-5.6 SolOpenAI82.8
70GPT-5.5 InstantOpenAI82.5
71Qwen3.5-35B-A3BAlibaba82.3
72Qwen3.5 FlashAlibaba82.3
73Qwen3.7 FlashAlibaba82.3
74Qwen3.7 PlusAlibaba81.8
75Claude Sonnet 4.5Anthropic81.7
76MiniMax M3MiniMax81.3
77o3OpenAI80.8
78Qwen3.7 FlashAlibaba80.8
79Claude Opus 4.5Anthropic80.7
80Qwen3-235B-A22BAlibaba80.0
81Qwen3.5 9BAlibaba79.0
82Qwen3.5 9BAlibaba78.9
83o4-miniOpenAI77.5
84DeepSeek-V3.2-ExpDeepSeek77.3
85GPT-5.5OpenAI77.3
86GPT-5.6 TerraOpenAI77.3
87Claude Sonnet 4Anthropic76.8
88Claude 3.7 SonnetAnthropic76.8
89Claude Opus 4.1Anthropic76.8
90DeepSeek-R1DeepSeek76.3
91Gemini 3.5 Flash-LiteGoogle75.8
92Gemma 4 31B ITGoogle75.8
93gpt-oss-120bOpenAI75.8
94Gemini 2.5 Pro PreviewGoogle75.8
95Grok-3 minixAI75.4
96GPT-5.4 nanoOpenAI75.3
97GPT-5.4OpenAI74.8
98Gemini 3.1 Flash-LiteGoogle74.2
99Claude Sonnet 4.5Anthropic73.7
100GPT 5 miniOpenAI73.3
101DeepSeek v4DeepSeek73.2
102Gemma 4 26B-A4B ITGoogle73.2
103GPT-5.2OpenAI73.2
104O3 MiniOpenAI73.2
105Claude Opus 4Anthropic72.7
106Qwen3 MaxAlibaba72.6
107seed-oss-36b-instruct71.5
108GLM-5.2Zhipu71.2
109Qwen3Alibaba70.7
110o1-2024-12-17OpenAI69.3
111DeepSeek R1DeepSeek69.2
112GPT-4.5OpenAI68.7
113DeepSeek-V3DeepSeek67.6
114Grok 3 BetaxAI67.6
115GPT 5 nanoOpenAI67.4
116Llama 4 Maverick InstructMeta67.0
117GPT 4.1OpenAI66.9
118GPT-5.1OpenAI66.7
119GPT 4.1 miniOpenAI65.8
120Claude Haiku 4.5Anthropic65.8
121Gemini 2.0 ProGoogle65.7
122QwQ-PlusAlibaba65.4
123QwQ 32BAlibaba65.3
124DeepSeek-R1-Distill-Qwen-32BDeepSeek64.1
125GPT-5.4 miniOpenAI64.1
126GPT-5.6 LunaOpenAI63.6
127Qwen3-30B-A3BAlibaba62.8
128o1-miniOpenAI60.9
129Gemini 2.0 Flash ThinkingGoogle60.6
130glm-4.7-flash60.5
131Qwen3 32BAlibaba59.9
132Mistral Medium 3Mistral59.5
133Qwen3-14BAlibaba58.6
134Qwen3 8BAlibaba56.8
135DeepSeek-V3DeepSeek56.5
136Qwen2.5-MaxAlibaba56.1
137Magistral Small 1.0Mistral56.1
138Phi-4Microsoft56.1
139Qwen3-30B-A3BAlibaba56.0
140DeepSeek-R1-Distill-Llama-70BDeepSeek55.7
141GPT-5.4 nanoOpenAI55.6
142Claude 3.5 SonnetAnthropic55.3
143Claude 3.5 SonnetAnthropic54.0
144Grok-2xAI53.8
145gpt-oss-20bOpenAI53.2
146Llama 4 Scout InstructMeta51.8
147Gemini 1.5 ProGoogle51.5
148Llama 3.1-405BMeta50.9
149Mistral Large 2Mistral50.2
150Qwen2.5-72BAlibaba49.1
151Mistral Small 3.2Mistral49.0
152GPT-4.1 nanoOpenAI48.9
153GPT 4oOpenAI48.7
154Qwen PlusAlibaba48.1
155Qwen3-4BAlibaba48.0
156Gemma 3 27BGoogle47.7
157Magistral Small 1.2Mistral47.6
158Qwen3 8BAlibaba47.5
159Llama 3.3 70BMeta47.4
160Claude 3 OpusAnthropic47.2
161GPT-4 TurboOpenAI46.6
162Mistral Small 3Mistral46.3
163Tülu 3 70BAi246.3
164Qwen2.5-32BAlibaba46.1
165qwen3-4b-instruct-250745.8
166glm-4.7-flash_none45.1
167DeepSeek-R1-Distill-Qwen-14BDeepSeek44.7
168Mistral Small 3.1Mistral44.7
169Llama 3.1-70BMeta44.2
170Gemini 1.5 FlashGoogle43.8
171WizardLM-2 8x22BMicrosoft43.4
172GPT-4 TurboOpenAI42.3
173Qwen TurboAlibaba41.8
174Llama 3.2 90BMeta41.0
175Qwen2-72BAlibaba40.8
176Claude 3 SonnetAnthropic40.6
177Llama 3-70BMeta40.6
178Gemma 3 12BGoogle39.5
179Mistral LargeMistral38.8
180Claude 3.5 HaikuAnthropic38.1
181GPT-4o miniOpenAI37.7
182Hermes 2 Theta Llama-3 70BNous Research37.5
183Gemma 2 27BGoogle36.5
184Claude 3 HaikuAnthropic36.3
185GPT-4OpenAI35.7
186Qwen2.5-7BAlibaba35.5
187Claude 2Anthropic34.7
188Qwen3-1.7BAlibaba34.3
189Mixtral 8x22BMistral34.1
190Gemini 1.0 ProGoogle34.0
191Eurus-2-7B-PRIMETsinghua University33.9
192DeepSeek-R1-Distill-Qwen-1.5BDeepSeek33.6
193Claude 2.1Anthropic33.0
194Gemini 1.5 Flash 8BGoogle33.0
195DBRXDatabricks32.9
196Yi-1.5-34B01.AI32.0
197Qwen1.5-32BAlibaba30.7
198GPT-4OpenAI30.6
199Mixtral 8x7BMistral30.2
200Mistral NeMoMistral29.9
201Qwen1.5-72BAlibaba28.8
202granite-4.0-micro28.3
203GPT-3.5 TurboOpenAI27.6
204phi-3-medium 14BMicrosoft27.6
205Gemma 2 9BGoogle27.5
206Ministral 8BMistral27.1
207Llama 3.1-8BMeta27.0
208Llama 2-70BMeta26.3
209Llama 3-8BMeta26.1
210Ministral 3BMistral25.2
211DeepSeek LLM 67BDeepSeek24.6
212granite-4.0-1b24.0
213Llama 3.2 1BMeta23.9
214Gemma 3 4BGoogle23.2
215Gemma 3 1BGoogle19.9
216Yi-34B01.AI14.7
217Mistral 7BMistral14.2
218granite-4.0-350m11.2
219DeepSeek-R1-05289.3

The benchmark's own page