MirrorCode
Coding · questions rotate
Coding tasks mirrored from recent real-world sources.
How far to trust it
Few models are measured on it, so its weight in the composite is small and its ranking is provisional.
Measured
Models scored
8
Spread between models
23.3
standard deviation, points
Measurement noise
8.4
published stderr
Weight
0.87
share of spread that is signal
The weight is (spread² − noise²) / spread². It is computed, never chosen. Most of this benchmark's spread is real, so it still separates models.
Every model scored
the median across configurations, so one heroic run cannot lead
| Rank | Model | Lab | Score |
|---|---|---|---|
| 1 | Claude Fable 5.1 | Anthropic | 73.3 |
| 2 | Claude Fable 5 | Anthropic | 63.9 |
| 3 | GPT-6 Astra | OpenAI | 46.7 |
| 4 | Claude Opus 4.7 | Anthropic | 31.1 |
| 5 | GPT-5.6 Sol | OpenAI | 20.0 |
| 6 | GPT-5.4 | OpenAI | 15.6 |
| 7 | GPT-5.5 | OpenAI | 10.0 |
| 8 | Gemini 3.1 Pro Preview | 8.9 |