OEIS Open Lite
Mathematics · fixed question set
Open problems from the Online Encyclopedia of Integer Sequences.
How far to trust it
🚨 Excluded from the composite by the weighting, not by hand. Its measured spread between models is barely larger than its own measurement error, so it carries almost no information about which model is better.
Measured
Models scored
5
Spread between models
5.4
standard deviation, points
Measurement noise
4.9
published stderr
Weight
0.18
share of spread that is signal
The weight is (spread² − noise²) / spread². It is computed, never chosen. This benchmark's spread is barely larger than its own error, so it says little about which model is better.
Every model scored
the median across configurations, so one heroic run cannot lead
| Rank | Model | Lab | Score |
|---|---|---|---|
| 1 | Claude Fable 5 | Anthropic | 44.0 |
| 2 | GPT-5.6 Sol | OpenAI | 43.0 |
| 3 | Claude Opus 4.8 | Anthropic | 39.0 |
| 4 | GPT-5.5 | OpenAI | 36.0 |
| 5 | Gemini 3.5 Flash | 29.0 |