modelbenchmark.io
All benchmarks

OEIS Open

Mathematics · fixed question set

The full OEIS open-problem set.

How far to trust it

Too few models are measured to place any model with confidence.

Measured

Models scored
3
Spread between models
standard deviation, points
Measurement noise
inferred from other benchmarks
Weight
excluded

The weight is (spread² − noise²) / spread². It is computed, never chosen. Most of this benchmark's spread is real, so it still separates models.

Every model scored

the median across configurations, so one heroic run cannot lead

View
RankModelLabScore
1Claude Opus 4.8Anthropic29.9
2GPT-5.5OpenAI26.2
3Gemini 3.5 FlashGoogle22.1

The benchmark's own page