modelbenchmark.io
All benchmarks

FrontierMath tiers 1-3

Mathematics · fixed question set

The lower three difficulty tiers of FrontierMath v2.

How far to trust it

Same private construction as FrontierMath v1. Reasoning effort moves this score more than any other benchmark on the site, so a figure with no effort setting attached says very little.

Measured

Models scored
85
Spread between models
26.0
standard deviation, points
Measurement noise
2.6
published stderr
Weight
0.99
share of spread that is signal

The weight is (spread² − noise²) / spread². It is computed, never chosen. Most of this benchmark's spread is real, so it still separates models.

Every model scored

the median across configurations, so one heroic run cannot lead

View
RankModelLabScore
1GPT-6 AstraOpenAI93.7
2Claude Fable 5.1Anthropic90.2
3GPT-5.6 SolOpenAI89.1
4GPT-5.5 ProOpenAI87.7
5Claude Fable 5Anthropic87.0
6GPT-5.6 TerraOpenAI86.0
7Claude Opus 5Anthropic85.6
8GPT-5.5OpenAI85.3
9GPT-5.4 ProOpenAI82.5
10Claude Opus 4.8Anthropic80.0
11GPT-5.4OpenAI78.6
12Qwen3.8 MaxAlibaba74.7
13GPT-5.2 ProOpenAI74.0
14Kimi K3Kimi72.2
15Gemini 3.7 FlashGoogle71.6
16Claude Opus 4.7Anthropic70.2
17GLM-5.3Zhipu68.8
18Gemini 3.8 FlashGoogle68.4
19GPT 5.2OpenAI67.4
20Claude Opus 4.6Anthropic66.0
21Grok 4.6xAI66.0
22Claude Sonnet 5Anthropic65.6
23Qwen3.8 Max (0902)Alibaba65.6
24DeepSeek V4 Pro 0813DeepSeek64.6
25Qwen3.7-MaxAlibaba64.6
26Gemini 3.5 FlashGoogle62.8
27GPT-5.6 LunaOpenAI61.8
28Gemini 3.1 Pro PreviewGoogle59.6
29Gemini 3.6 FlashGoogle59.0
30DeepSeek V4 Flash 0731DeepSeek57.5
31Grok 4.5xAI57.2
32Kimi K2.6Kimi57.2
33GLM-5.2Zhipu57.0
34GLM-5.3-FlashZhipu55.8
35GPT-5 ProOpenAI55.8
36Kimi K2.7 CodeKimi54.0
37Gemini 3 Flash PreviewGoogle51.2
38Inkling SmallThinking Machines46.3
39DeepSeek v4DeepSeek45.3
40Grok 4.20xAI44.9
41Grok 4.3 BetaxAI42.8
42GLM-5.2Zhipu42.5
43GPT-5.6 LunaOpenAI39.6
44Qwen3.6 PlusAlibaba38.2
45GPT-5.4 miniOpenAI37.9
46GPT 5OpenAI37.2
47Qwen3.6 27BAlibaba35.1
48Claude Opus 4.5Anthropic34.4
49Qwen3.7 PlusAlibaba34.4
50Qwen3.6 27BAlibaba34.0
51InklingThinking Machines33.3
52GPT-5.4 nanoOpenAI32.6
53Qwen3.6 PlusAlibaba32.3
54Qwen3.5 397B-A17BAlibaba31.2
55GLM-5.1Zhipu30.9
56Qwen3.5 397B-A17BAlibaba29.5
57GPT 5 miniOpenAI29.4
58o3OpenAI27.5
59o4-miniOpenAI27.0
60GPT-5.5 InstantOpenAI26.3
61Gemini 3.5 Flash-LiteGoogle26.0
62Gemini 2.5 ProGoogle24.6
63Claude Sonnet 4.5Anthropic23.9
64Gemini 3.1 Flash-LiteGoogle22.5
65Qwen3.6 FlashAlibaba22.5
66Qwen3.6 35B-A3BAlibaba20.4
67Qwen3.7 FlashAlibaba19.3
68Qwen3.7 FlashAlibaba19.3
69Qwen3 MaxAlibaba18.9
70Qwen3.5 FlashAlibaba18.2
71Qwen3.6 35B-A3BAlibaba17.5
72GPT-5.4 miniOpenAI17.2
73Qwen3.6 FlashAlibaba17.2
74Claude Opus 4.1Anthropic12.6
75GPT 5 nanoOpenAI11.9
76o1-2024-12-17OpenAI11.1
77O3 MiniOpenAI11.0
78Qwen3.5 FlashAlibaba9.5
79GPT 4.1 miniOpenAI6.7
80GPT 4.1OpenAI6.0
81GPT-5.4 nanoOpenAI4.6
82GPT-4 TurboOpenAI0.7
83GPT-4o miniOpenAI0.7
84GPT 4oOpenAI0.3
85GPT-3.5 TurboOpenAI0.0

The benchmark's own page