modelbenchmark.io
All benchmarks

SimpleQA Verified

Knowledge · fixed question set

Short factual questions with one verifiable answer.

How far to trust it

Measures recall and the willingness to say 'I do not know'. A model that guesses confidently scores worse here than one that declines, which is the opposite of most benchmarks.

Measured

Models scored
72
Spread between models
18.1
standard deviation, points
Measurement noise
1.5
published stderr
Weight
0.99
share of spread that is signal

The weight is (spread² − noise²) / spread². It is computed, never chosen. Most of this benchmark's spread is real, so it still separates models.

Every model scored

the median across configurations, so one heroic run cannot lead

View
RankModelLabScore
1GPT-6 AstraOpenAI75.6
2Gemini 3.1 Pro PreviewGoogle73.5
3Claude Fable 5.1Anthropic70.8
4Claude Fable 5Anthropic70.7
5Gemini 3.8 FlashGoogle69.7
6GPT-5.6 SolOpenAI69.7
7Gemini 3.7 FlashGoogle69.2
8Gemini 3 Flash PreviewGoogle66.8
9Gemini 3.5 FlashGoogle66.2
10Gemini 3.6 FlashGoogle66.2
11GPT-5.5OpenAI63.0
12Muse Spark 1.2Meta60.3
13Claude Opus 5Anthropic59.9
14Muse Spark 1.1Meta57.8
15Qwen3.7-MaxAlibaba55.8
16Claude Opus 4.8Anthropic53.0
17DeepSeek V4 Pro 0813DeepSeek52.9
18Qwen3.6 Max PreviewAlibaba52.0
19Claude Opus 4.7Anthropic51.7
20Kimi K3Kimi50.6
21GPT 5OpenAI50.1
22o3OpenAI49.4
23Grok 4.6xAI49.1
24Qwen3 MaxAlibaba48.8
25Grok 4.5xAI48.3
26GPT 5.1OpenAI48.0
27Qwen3.8 Max (0902)Alibaba47.3
28Claude Opus 4.6Anthropic47.0
29DeepSeek v4DeepSeek47.0
30GPT-5.4 ProOpenAI46.3
31Qwen3.8 MaxAlibaba45.8
32Claude Opus 4.5Anthropic45.7
33GPT-5.4OpenAI45.1
34Qwen3.6 PlusAlibaba44.1
35GPT-5.6 TerraOpenAI43.2
36o1-2024-12-17OpenAI41.1
37GLM-5.3Zhipu41.0
38GPT-5.6 LunaOpenAI41.0
39Qwen3-235B-A22BAlibaba40.4
40InklingThinking Machines40.3
41Kimi K2.7 CodeKimi36.5
42Kimi K2.6Kimi34.9
43Kimi K2.5Kimi34.3
44GLM-5.2Zhipu34.2
45Claude Sonnet 4.6Anthropic34.1
46GLM-5.1Zhipu34.0
47DeepSeek V4 Flash 0731DeepSeek33.6
48GPT 5.2OpenAI33.5
49Claude Sonnet 5Anthropic33.3
50Grok 4.3 BetaxAI33.2
51GLM-4.7Zhipu32.2
52GPT 4.1OpenAI31.1
53Claude Sonnet 4.5Anthropic30.7
54Grok 4.20xAI30.2
55GPT-5.4 miniOpenAI29.4
56GPT 4oOpenAI26.0
57Qwen3.5 PlusAlibaba25.4
58Claude Sonnet 4.5Anthropic23.7
59GPT 5 miniOpenAI21.6
60Qwen3.5 FlashAlibaba20.3
61o4-miniOpenAI19.2
62Inkling SmallThinking Machines19.1
63Qwen3.6 FlashAlibaba15.9
64O3 MiniOpenAI15.3
65Claude Haiku 4.5Anthropic12.9
66GPT 4.1 miniOpenAI12.7
67Claude 3 OpusAnthropic12.6
68GPT-5.4 nanoOpenAI11.7
69GPT 5 nanoOpenAI11.7
70Gemma 4 31B ITGoogle10.4
71GPT-4o miniOpenAI8.3
72GPT-4.1 nanoOpenAI6.0

The benchmark's own page