OEIS Open
Mathematics · fixed question set
The full OEIS open-problem set.
How far to trust it
Too few models are measured to place any model with confidence.
Measured
Models scored
3
Spread between models
—
standard deviation, points
Measurement noise
—
inferred from other benchmarks
Weight
—
excluded
The weight is (spread² − noise²) / spread². It is computed, never chosen. Most of this benchmark's spread is real, so it still separates models.
Every model scored
the median across configurations, so one heroic run cannot lead
| Rank | Model | Lab | Score |
|---|---|---|---|
| 1 | Claude Opus 4.8 | Anthropic | 29.9 |
| 2 | GPT-5.5 | OpenAI | 26.2 |
| 3 | Gemini 3.5 Flash | 22.1 |