How the score is built
Benchmarks
A benchmark earns its place in the composite by discriminating. We measure two things for each one: how far apart the models are on it, and how far apart two measurements of the same model are. The weight is the share of the first that survives the second.
Weight = (spread² − noise²) / spread²A benchmark every model passes has almost no spread left once its own measurement error is subtracted, so it stops counting on its own. No weight on this page was chosen by hand.
Sources
| Source | Who ran it | Where |
|---|---|---|
| Epoch AI | independent runs | epoch.ai/benchmarks |
| SWE-bench | benchmark authors | www.swebench.com |
| LiveBench | benchmark authors | livebench.ai |
| Aider | benchmark authors | aider.chat/docs/leaderboards |
Epoch AI data is used under CC-BY.
In the composite
MATH level 5
w 1.00Mathematics
The hardest tier of the MATH competition set.
87 models · spread 30.7 · noise 0.9
SimpleQA Verified
w 0.99Knowledge
Short factual questions with one verifiable answer.
72 models · spread 18.1 · noise 1.5
FrontierMath tiers 1-3
w 0.99Mathematics
The lower three difficulty tiers of FrontierMath v2.
85 models · spread 26.0 · noise 2.6
SWE-bench Verified
w 0.99Coding
500 real GitHub issues from Python repositories, human-checked to be solvable. The model must produce a patch that passes the repository's own tests.
84 models · spread 20.4 · noise 2.0
FrontierMath v1
w 0.99Mathematics
Research-level mathematics problems, most written by working mathematicians for this benchmark.
81 models · spread 23.5 · noise 2.6
GPQA Diamond
w 0.99Reasoning
448 graduate-level questions in biology, physics and chemistry, written so that a skilled non-expert with web access still scores near 34%.
219 models · spread 22.8 · noise 2.5
Aider polyglot
w 0.99Coding
225 Exercism exercises across C++, Go, Java, JavaScript, Python and Rust, edited in place through the aider tool.
60 models · spread 23.6 · noise 2.7 (inferred)
OTIS Mock AIME
w 0.98Mathematics · rotating
Mock AIME problems written for the OTIS programme, answered as an integer from 0 to 999.
200 models · spread 35.1 · noise 4.9
EBR-bench
w 0.98Reasoning
Evidence-based reasoning tasks.
21 models · spread 19.0 · noise 2.8
FrontierMath tier 4
w 0.95Mathematics
The hardest FrontierMath tier: problems aimed at research mathematicians.
64 models · spread 16.3 · noise 3.5
FrontierMath tier 4 v2
w 0.94Mathematics
The second version of the hardest FrontierMath tier.
58 models · spread 27.3 · noise 6.5
Chess Puzzles
w 0.94Reasoning · rotating
Tactical chess positions with one correct continuation.
157 models · spread 15.8 · noise 3.8
Mystery Game Puzzles
w 0.93Reasoning · rotating
Multi-step deduction inside a game the model must reason about from its rules.
86 models · spread 13.8 · noise 3.8
MirrorCode
w 0.87Coding · rotating
Coding tasks mirrored from recent real-world sources.
8 models · spread 23.3 · noise 8.4
LiveBench
w 0.72General · rotating
23 subtasks over maths, coding, reasoning, language, data analysis and instruction following. Questions are refreshed from recent sources.
57 models · spread 5.1 · noise 2.7 (inferred)
OEIS Open Lite
w 0.18Mathematics
Open problems from the Online Encyclopedia of Integer Sequences.
5 models · spread 5.4 · noise 4.9