LiveBench
General · questions rotate
23 subtasks over maths, coding, reasoning, language, data analysis and instruction following. Questions are refreshed from recent sources.
How far to trust it
Contamination-resistant by design, because questions rotate. 🚨 The 23 subtasks are averaged into one figure here. They are not interchangeable: table reformatting scores near 98 where TypeScript scores near 20. The average compresses models together, so it separates them less than a single hard benchmark does.
Measured
Models scored
57
Spread between models
5.1
standard deviation, points
Measurement noise
2.7
inferred from other benchmarks
Weight
0.72
share of spread that is signal
The weight is (spread² − noise²) / spread². It is computed, never chosen. Most of this benchmark's spread is real, so it still separates models.
Every model scored
the median across configurations, so one heroic run cannot lead