modelbenchmark.io
All benchmarks

FrontierMath tier 4 v2

Mathematics · fixed question set

The second version of the hardest FrontierMath tier.

How far to trust it

As tier 4. The v2 problems are newer, so contamination is lower still, and the two versions are not directly comparable.

Measured

Models scored
58
Spread between models
27.3
standard deviation, points
Measurement noise
6.5
published stderr
Weight
0.94
share of spread that is signal

The weight is (spread² − noise²) / spread². It is computed, never chosen. Most of this benchmark's spread is real, so it still separates models.

Every model scored

the median across configurations, so one heroic run cannot lead

View
RankModelLabScore
1GPT-6 AstraOpenAI97.6
2Claude Fable 5Anthropic90.2
3Claude Fable 5.1Anthropic87.8
4GPT-5.6 SolOpenAI82.9
5GPT-6 AstraOpenAI82.9
6GPT-5.6 SolOpenAI80.5
7GPT-5.5 ProOpenAI78.0
8AI Co-MathematicianGoogle75.6
9Claude Opus 5Anthropic73.2
10GPT-5.5OpenAI72.5
11GPT-5.6 TerraOpenAI70.7
12GPT-5.6 LunaOpenAI61.0
13GPT-5.4 ProOpenAI58.5
14Claude Opus 4.8Anthropic56.1
15GPT-5.4OpenAI49.0
16Qwen3.8 MaxAlibaba46.3
17GPT-5.2 ProOpenAI46.0
18Kimi K3Kimi39.0
19Gemini 3.7 FlashGoogle36.6
20Qwen3.7-MaxAlibaba34.1
21Qwen3.8 Max (0902)Alibaba34.1
22Claude Opus 4.7Anthropic31.7
23Grok 4.6xAI31.7
24GPT 5.2OpenAI31.7
25Claude Sonnet 5Anthropic29.3
26GLM-5.2Zhipu29.3
27GLM-5.3Zhipu29.3
28Claude Opus 4.6Anthropic26.8
29DeepSeek V4 Pro 0813DeepSeek26.8
30Gemini 3.1 Pro PreviewGoogle26.8
31Gemini 3.5 FlashGoogle26.8
32Kimi K2.6Kimi25.6
33DeepSeek V4 Flash 0731DeepSeek24.4
34Grok 4.5xAI24.4
35Gemini 3.6 FlashGoogle21.9
36Gemini 3.8 FlashGoogle21.9
37GPT 5OpenAI21.9
38GPT-5 ProOpenAI19.5
39Gemini 3 Flash PreviewGoogle17.1
40GLM-5.3-FlashZhipu17.1
41Grok 4.20xAI17.1
42Inkling SmallThinking Machines17.1
43Grok 4.3 BetaxAI14.6
44GPT-5.4 nanoOpenAI12.2
45GPT 5 miniOpenAI12.2
46Kimi K2.7 CodeKimi12.2
47GPT-5.4 miniOpenAI9.8
48Claude Opus 4.5Anthropic4.9
49InklingThinking Machines4.9
50o4-miniOpenAI4.9
51Claude Opus 4.1Anthropic2.4
52Claude Sonnet 4.5Anthropic2.4
53DeepSeek v4DeepSeek2.4
54GPT-5.5 InstantOpenAI2.4
55GPT 5 nanoOpenAI2.4
56Gemini 2.5 ProGoogle0.0
57Gemini 3.5 Flash-LiteGoogle0.0
58O3 MiniOpenAI0.0

The benchmark's own page