modelbenchmark.io
All benchmarks

OEIS Open Lite

Mathematics · fixed question set

Open problems from the Online Encyclopedia of Integer Sequences.

How far to trust it

🚨 Excluded from the composite by the weighting, not by hand. Its measured spread between models is barely larger than its own measurement error, so it carries almost no information about which model is better.

Measured

Models scored
5
Spread between models
5.4
standard deviation, points
Measurement noise
4.9
published stderr
Weight
0.18
share of spread that is signal

The weight is (spread² − noise²) / spread². It is computed, never chosen. This benchmark's spread is barely larger than its own error, so it says little about which model is better.

Every model scored

the median across configurations, so one heroic run cannot lead

View
RankModelLabScore
1Claude Fable 5Anthropic44.0
2GPT-5.6 SolOpenAI43.0
3Claude Opus 4.8Anthropic39.0
4GPT-5.5OpenAI36.0
5Gemini 3.5 FlashGoogle29.0

The benchmark's own page