modelbenchmark.io
All benchmarks

MirrorCode

Coding · questions rotate

Coding tasks mirrored from recent real-world sources.

How far to trust it

Few models are measured on it, so its weight in the composite is small and its ranking is provisional.

Measured

Models scored
8
Spread between models
23.3
standard deviation, points
Measurement noise
8.4
published stderr
Weight
0.87
share of spread that is signal

The weight is (spread² − noise²) / spread². It is computed, never chosen. Most of this benchmark's spread is real, so it still separates models.

Every model scored

the median across configurations, so one heroic run cannot lead

View
RankModelLabScore
1Claude Fable 5.1Anthropic73.3
2Claude Fable 5Anthropic63.9
3GPT-6 AstraOpenAI46.7
4Claude Opus 4.7Anthropic31.1
5GPT-5.6 SolOpenAI20.0
6GPT-5.4OpenAI15.6
7GPT-5.5OpenAI10.0
8Gemini 3.1 Pro PreviewGoogle8.9

The benchmark's own page