modelbenchmark.io
All benchmarks

LiveBench

General · questions rotate

23 subtasks over maths, coding, reasoning, language, data analysis and instruction following. Questions are refreshed from recent sources.

How far to trust it

Contamination-resistant by design, because questions rotate. 🚨 The 23 subtasks are averaged into one figure here. They are not interchangeable: table reformatting scores near 98 where TypeScript scores near 20. The average compresses models together, so it separates them less than a single hard benchmark does.

Measured

Models scored
57
Spread between models
5.1
standard deviation, points
Measurement noise
2.7
inferred from other benchmarks
Weight
0.72
share of spread that is signal

The weight is (spread² − noise²) / spread². It is computed, never chosen. Most of this benchmark's spread is real, so it still separates models.

Every model scored

the median across configurations, so one heroic run cannot lead

View
RankModelLabScore
1Claude Fable 5.1Anthropic83.8
2Claude Fable 5Anthropic83.4
3GPT-6 AstraOpenAI83.0
4muse-spark-1.3-xhigh82.4
5GPT-5.6 SolOpenAI81.6
6deepseek-v4.1-flash-max81.4
7GPT-5.5OpenAI80.8
8Claude Opus 5Anthropic80.5
9Gemini 3.7 FlashGoogle79.9
10smaug-agentic79.7
11Qwen3.8 MaxAlibaba79.5
12Kimi K3Kimi79.5
13Grok 4.6xAI79.0
14Muse Spark 1.2Meta78.9
15GPT-5.4OpenAI78.8
16GPT-5.6 TerraOpenAI78.6
17DeepSeek V4 Pro 0813DeepSeek78.2
18smaug-flash78.0
19Gemini 3.1 Pro PreviewGoogle78.0
20deepseek-v4-flash-vision-exp77.7
21smaug-mini77.7
22Gemini 3.8 FlashGoogle77.5
23qwen3.8-flash-next77.3
24Claude Opus 4.8Anthropic77.1
25Grok 4.5xAI77.1
26Claude Opus 4.7Anthropic77.0
27GLM-5.3Zhipu76.6
28Claude Sonnet 5Anthropic76.6
29Muse Spark 1.1Meta76.0
30qwen3.8-27b75.8
31Gemini 3.5 FlashGoogle75.4
32GPT 5.2OpenAI75.2
33Claude Opus 4.6Anthropic75.1
34DeepSeek V4 Flash 0731DeepSeek74.8
35Gemini 3.6 FlashGoogle74.5
36Qwen3.7-MaxAlibaba74.1
37GPT 5.2 CodexOpenAI74.0
38GPT-5.6 LunaOpenAI73.7
39Claude Sonnet 4.6Anthropic73.4
40GLM-5.2Zhipu73.4
41Claude Opus 4.5Anthropic72.9
42InklingThinking Machines72.9
43deepseek-v4-pro72.6
44GLM-5.3-FlashZhipu71.1
45Kimi K2.6Kimi70.9
46GPT-5.4 nanoOpenAI70.8
47ox-alpha-max69.2
48Qwen3.6 PlusAlibaba69.0
49Kimi K2.7 CodeKimi68.8
50nemotron-3-ultra-550b-a55b68.7
51grok-build-0.168.6
52MiniMax M3MiniMax67.5
53GPT-5.4 miniOpenAI66.6
54deepseek-v4-flash66.1
55Qwen3.6 27BAlibaba64.2
56Gemini 3.5 Flash-LiteGoogle63.8
57grok-4.363.3

The benchmark's own page