modelbenchmark.io
All benchmarks

MATH level 5

Mathematics · fixed question set

The hardest tier of the MATH competition set.

How far to trust it

Old, public and thoroughly in the training data of every recent model. It still separates models because the tier is hard, but read a high score as evidence of competence, never of novelty.

Measured

Models scored
87
Spread between models
30.7
standard deviation, points
Measurement noise
0.9
published stderr
Weight
1.00
share of spread that is signal

The weight is (spread² − noise²) / spread². It is computed, never chosen. Most of this benchmark's spread is real, so it still separates models.

Every model scored

the median across configurations, so one heroic run cannot lead

View
RankModelLabScore
1GPT 5OpenAI98.0
2o4-miniOpenAI97.8
3o3OpenAI97.8
4Claude Sonnet 4.5Anthropic97.7
5GPT 5 miniOpenAI97.3
6Qwen3 MaxAlibaba97.1
7DeepSeek-R1DeepSeek96.6
8Gemini 2.5 Pro PreviewGoogle95.9
9O3 MiniOpenAI95.8
10Gemini 2.5 Pro ExpGoogle95.6
11GPT 5 nanoOpenAI95.1
12DeepSeek R1DeepSeek93.0
13Claude Haiku 4.5Anthropic91.6
14o1-2024-12-17OpenAI90.3
15DeepSeek-R1-Distill-Llama-70BDeepSeek89.9
16Grok-3 minixAI89.5
17Grok 3 BetaxAI88.8
18Claude 3.7 SonnetAnthropic88.1
19GPT 4.1 miniOpenAI87.3
20DeepSeek-R1-Distill-Qwen-14BDeepSeek87.1
21o1-miniOpenAI86.7
22Claude Opus 4Anthropic85.0
23Claude Sonnet 4Anthropic84.4
24Gemini 2.0 ProGoogle83.5
25GPT 4.1OpenAI83.0
26Gemini 2.0 Flash ThinkingGoogle82.2
27Mistral Medium 3Mistral81.6
28GPT-4.5OpenAI78.6
29DeepSeek-V3DeepSeek75.5
30Gemma 3 27BGoogle74.0
31Llama 4 Maverick InstructMeta73.0
32GPT-4.1 nanoOpenAI70.0
33Qwen3Alibaba68.9
34Qwen2.5-MaxAlibaba67.2
35Qwen PlusAlibaba65.3
36Phi-4Microsoft64.9
37DeepSeek-V3DeepSeek64.8
38Grok-2xAI63.5
39Qwen2.5-72BAlibaba63.2
40Llama 4 Scout InstructMeta62.3
41Claude 3.5 SonnetAnthropic57.0
42Qwen TurboAlibaba56.2
43Qwen2.5-32BAlibaba56.1
44Gemini 1.5 ProGoogle55.6
45GPT-4o miniOpenAI52.6
46Claude 3.5 SonnetAnthropic51.7
47GPT 4oOpenAI51.4
48Llama 3.1-405BMeta49.8
49Mistral Large 2Mistral47.5
50Mistral Small 3.1Mistral46.8
51GPT-4 TurboOpenAI46.7
52Claude 3.5 HaikuAnthropic46.4
53Mistral Small 3Mistral44.8
54Gemini 1.5 FlashGoogle43.5
55Tülu 3 70BAi242.7
56Llama 3.3 70BMeta41.6
57Llama 3.2 90BMeta39.4
58Qwen2-72BAlibaba39.1
59GPT-4 TurboOpenAI37.7
60Claude 3 OpusAnthropic37.5
61Llama 3.1-70BMeta36.7
62Gemma 2 27BGoogle27.9
63WizardLM-2 8x22BMicrosoft25.7
64Yi-1.5-34B01.AI25.5
65Mistral LargeMistral24.5
66Mixtral 8x22BMistral24.2
67GPT-4OpenAI23.0
68Llama 3.1-8BMeta22.9
69Hermes 2 Theta Llama-3 70BNous Research22.7
70Llama 3-70BMeta22.6
71Gemma 2 9BGoogle21.0
72Claude 3 SonnetAnthropic18.2
73phi-3-medium 14BMicrosoft17.6
74Ministral 8BMistral14.9
75Claude 3 HaikuAnthropic14.9
76Ministral 3BMistral14.4
77GPT-3.5 TurboOpenAI13.8
78Claude 2Anthropic11.7
79DBRXDatabricks11.7
80Gemini 1.0 ProGoogle11.2
81Mistral NeMoMistral10.8
82Mixtral 8x7BMistral9.6
83DeepSeek LLM 67BDeepSeek6.4
84Llama 3-8BMeta6.1
85Yi-34B01.AI5.2
86Mistral 7BMistral3.6
87Llama 2-70BMeta3.3

The benchmark's own page