modelbenchmark.io
All benchmarks

FrontierMath v1

Mathematics · fixed question set

Research-level mathematics problems, most written by working mathematicians for this benchmark.

How far to trust it

Almost all problems are held private, so contamination is very low. The trade-off is that you cannot inspect what was asked. Ten problems are public; every figure here is from the private set.

Measured

Models scored
81
Spread between models
23.5
standard deviation, points
Measurement noise
2.6
published stderr
Weight
0.99
share of spread that is signal

The weight is (spread² − noise²) / spread². It is computed, never chosen. Most of this benchmark's spread is real, so it still separates models.

Every model scored

the median across configurations, so one heroic run cannot lead

View
RankModelLabScore
1Claude Fable 5Anthropic100.0
2Claude Opus 4.7Anthropic66.9
3Claude Opus 4.6Anthropic64.8
4Kimi K2.6Kimi64.5
5GPT-5.4OpenAI63.8
6Gemini 3.1 Pro PreviewGoogle62.9
7GPT-5 ProOpenAI60.0
8Gemini 3.5 FlashGoogle59.5
9Gemini 3 Pro PreviewGoogle58.8
10GLM-5.1Zhipu56.7
11Claude Sonnet 4.6Anthropic56.2
12GPT-5.5OpenAI51.7
13GPT-5.5 Pro pre-releaseOpenAI51.7
14GPT-5.4 ProOpenAI50.0
15Gemini 3 Flash PreviewGoogle47.8
16GPT 5OpenAI47.4
17Claude Opus 4.8Anthropic47.2
18GPT 5.2OpenAI44.6
19GPT-5.4 nanoOpenAI42.9
20GPT-5.4 miniOpenAI39.1
21Muse SparkMeta39.0
22Claude Opus 4.6Anthropic38.3
23Qwen3.6 PlusAlibaba38.1
24Qwen3.6 Max PreviewAlibaba36.5
25Qwen3.5 PlusAlibaba35.5
26GPT 5 miniOpenAI29.4
27Gemini 2.5 Deep ThinkGoogle29.0
28Kimi K2.5Kimi27.9
29Kimi K2 InstructKimi27.1
30Gemini 2.5 ProGoogle27.1
31GPT 5.1OpenAI26.9
32Claude Opus 4.5Anthropic25.4
33o4-miniOpenAI22.4
34DeepSeek-V3.2-ExpDeepSeek22.1
35Claude Opus 4.5Anthropic20.3
36Gemini 2.5 Pro PreviewGoogle20.2
37Grok 4xAI19.7
38O3 MiniOpenAI17.9
39Qwen3.6 FlashAlibaba15.2
40Claude Sonnet 4.5Anthropic14.4
41Qwen3-235B-A22BAlibaba14.2
42GPT 5 nanoOpenAI11.4
43o3OpenAI10.9
44GLM 5Zhipu9.3
45Qwen3.5 FlashAlibaba8.1
46GPT 4.1 miniOpenAI7.2
47Gemini 2.5 FlashGoogle4.8
48Claude Sonnet 4.5Anthropic4.7
49o1-2024-12-17OpenAI4.7
50Claude Haiku 4.5Anthropic3.5
51Claude Opus 4Anthropic3.2
52Claude Opus 4.1Anthropic2.9
53GPT 4.1OpenAI2.8
54GLM-4.7Zhipu2.4
55Grok-3 minixAI2.2
56GPT-5.1OpenAI2.1
57Claude Sonnet 4Anthropic2.1
58GLM 4.6Zhipu1.9
59Grok 3 BetaxAI1.9
60Claude 3.7 SonnetAnthropic1.6
61Claude 3.5 SonnetAnthropic1.0
62DeepSeek-V3DeepSeek0.9
63Gemini 2.0 Flash ThinkingGoogle0.9
64Qwen PlusAlibaba0.9
65o1-miniOpenAI0.8
66Claude 3.5 SonnetAnthropic0.5
67GPT-4.1 nanoOpenAI0.5
68Qwen2.5-MaxAlibaba0.5
69Grok-2xAI0.3
70Llama 4 Maverick InstructMeta0.3
71Mistral Medium 3Mistral0.2
72Claude 3.5 HaikuAnthropic0.2
73GPT 4oOpenAI0.2
74Mistral Large 2Mistral0.2
75DeepSeek-V3DeepSeek0.0
76Gemini 1.5 FlashGoogle0.0
77Gemini 2.0 ProGoogle0.0
78GLM 4.5Zhipu0.0
79Llama 4 Scout InstructMeta0.0
80Magistral Small 1.0Mistral0.0
81Qwen3Alibaba0.0

The benchmark's own page