MATH level 5
Mathematics · fixed question set
The hardest tier of the MATH competition set.
How far to trust it
Old, public and thoroughly in the training data of every recent model. It still separates models because the tier is hard, but read a high score as evidence of competence, never of novelty.
Measured
Models scored
87
Spread between models
30.7
standard deviation, points
Measurement noise
0.9
published stderr
Weight
1.00
share of spread that is signal
The weight is (spread² − noise²) / spread². It is computed, never chosen. Most of this benchmark's spread is real, so it still separates models.
Every model scored
the median across configurations, so one heroic run cannot lead
| Rank | Model | Lab | Score |
|---|---|---|---|
| 1 | GPT 5 | OpenAI | 98.0 |
| 2 | o4-mini | OpenAI | 97.8 |
| 3 | o3 | OpenAI | 97.8 |
| 4 | Claude Sonnet 4.5 | Anthropic | 97.7 |
| 5 | GPT 5 mini | OpenAI | 97.3 |
| 6 | Qwen3 Max | Alibaba | 97.1 |
| 7 | DeepSeek-R1 | DeepSeek | 96.6 |
| 8 | Gemini 2.5 Pro Preview | 95.9 | |
| 9 | O3 Mini | OpenAI | 95.8 |
| 10 | Gemini 2.5 Pro Exp | 95.6 | |
| 11 | GPT 5 nano | OpenAI | 95.1 |
| 12 | DeepSeek R1 | DeepSeek | 93.0 |
| 13 | Claude Haiku 4.5 | Anthropic | 91.6 |
| 14 | o1-2024-12-17 | OpenAI | 90.3 |
| 15 | DeepSeek-R1-Distill-Llama-70B | DeepSeek | 89.9 |
| 16 | Grok-3 mini | xAI | 89.5 |
| 17 | Grok 3 Beta | xAI | 88.8 |
| 18 | Claude 3.7 Sonnet | Anthropic | 88.1 |
| 19 | GPT 4.1 mini | OpenAI | 87.3 |
| 20 | DeepSeek-R1-Distill-Qwen-14B | DeepSeek | 87.1 |
| 21 | o1-mini | OpenAI | 86.7 |
| 22 | Claude Opus 4 | Anthropic | 85.0 |
| 23 | Claude Sonnet 4 | Anthropic | 84.4 |
| 24 | Gemini 2.0 Pro | 83.5 | |
| 25 | GPT 4.1 | OpenAI | 83.0 |
| 26 | Gemini 2.0 Flash Thinking | 82.2 | |
| 27 | Mistral Medium 3 | Mistral | 81.6 |
| 28 | GPT-4.5 | OpenAI | 78.6 |
| 29 | DeepSeek-V3 | DeepSeek | 75.5 |
| 30 | Gemma 3 27B | 74.0 | |
| 31 | Llama 4 Maverick Instruct | Meta | 73.0 |
| 32 | GPT-4.1 nano | OpenAI | 70.0 |
| 33 | Qwen3 | Alibaba | 68.9 |
| 34 | Qwen2.5-Max | Alibaba | 67.2 |
| 35 | Qwen Plus | Alibaba | 65.3 |
| 36 | Phi-4 | Microsoft | 64.9 |
| 37 | DeepSeek-V3 | DeepSeek | 64.8 |
| 38 | Grok-2 | xAI | 63.5 |
| 39 | Qwen2.5-72B | Alibaba | 63.2 |
| 40 | Llama 4 Scout Instruct | Meta | 62.3 |
| 41 | Claude 3.5 Sonnet | Anthropic | 57.0 |
| 42 | Qwen Turbo | Alibaba | 56.2 |
| 43 | Qwen2.5-32B | Alibaba | 56.1 |
| 44 | Gemini 1.5 Pro | 55.6 | |
| 45 | GPT-4o mini | OpenAI | 52.6 |
| 46 | Claude 3.5 Sonnet | Anthropic | 51.7 |
| 47 | GPT 4o | OpenAI | 51.4 |
| 48 | Llama 3.1-405B | Meta | 49.8 |
| 49 | Mistral Large 2 | Mistral | 47.5 |
| 50 | Mistral Small 3.1 | Mistral | 46.8 |
| 51 | GPT-4 Turbo | OpenAI | 46.7 |
| 52 | Claude 3.5 Haiku | Anthropic | 46.4 |
| 53 | Mistral Small 3 | Mistral | 44.8 |
| 54 | Gemini 1.5 Flash | 43.5 | |
| 55 | Tülu 3 70B | Ai2 | 42.7 |
| 56 | Llama 3.3 70B | Meta | 41.6 |
| 57 | Llama 3.2 90B | Meta | 39.4 |
| 58 | Qwen2-72B | Alibaba | 39.1 |
| 59 | GPT-4 Turbo | OpenAI | 37.7 |
| 60 | Claude 3 Opus | Anthropic | 37.5 |
| 61 | Llama 3.1-70B | Meta | 36.7 |
| 62 | Gemma 2 27B | 27.9 | |
| 63 | WizardLM-2 8x22B | Microsoft | 25.7 |
| 64 | Yi-1.5-34B | 01.AI | 25.5 |
| 65 | Mistral Large | Mistral | 24.5 |
| 66 | Mixtral 8x22B | Mistral | 24.2 |
| 67 | GPT-4 | OpenAI | 23.0 |
| 68 | Llama 3.1-8B | Meta | 22.9 |
| 69 | Hermes 2 Theta Llama-3 70B | Nous Research | 22.7 |
| 70 | Llama 3-70B | Meta | 22.6 |
| 71 | Gemma 2 9B | 21.0 | |
| 72 | Claude 3 Sonnet | Anthropic | 18.2 |
| 73 | phi-3-medium 14B | Microsoft | 17.6 |
| 74 | Ministral 8B | Mistral | 14.9 |
| 75 | Claude 3 Haiku | Anthropic | 14.9 |
| 76 | Ministral 3B | Mistral | 14.4 |
| 77 | GPT-3.5 Turbo | OpenAI | 13.8 |
| 78 | Claude 2 | Anthropic | 11.7 |
| 79 | DBRX | Databricks | 11.7 |
| 80 | Gemini 1.0 Pro | 11.2 | |
| 81 | Mistral NeMo | Mistral | 10.8 |
| 82 | Mixtral 8x7B | Mistral | 9.6 |
| 83 | DeepSeek LLM 67B | DeepSeek | 6.4 |
| 84 | Llama 3-8B | Meta | 6.1 |
| 85 | Yi-34B | 01.AI | 5.2 |
| 86 | Mistral 7B | Mistral | 3.6 |
| 87 | Llama 2-70B | Meta | 3.3 |