modelbenchmark.io
All benchmarks

EBR-bench

Reasoning · fixed question set

Evidence-based reasoning tasks.

How far to trust it

Only 19 models are measured, so most models cannot be placed on it at all. Its weight is real but its coverage is thin.

Measured

Models scored
21
Spread between models
19.0
standard deviation, points
Measurement noise
2.8
published stderr
Weight
0.98
share of spread that is signal

The weight is (spread² − noise²) / spread². It is computed, never chosen. Most of this benchmark's spread is real, so it still separates models.

Every model scored

the median across configurations, so one heroic run cannot lead

View
RankModelLabScore
1GPT-6 AstraOpenAI76.2
2Claude Fable 5.1Anthropic57.1
3Claude Opus 5Anthropic45.7
4GPT-5.6 SolOpenAI44.8
5Claude Fable 5Anthropic39.5
6GPT-5.5OpenAI34.3
7Grok 4.6xAI30.5
8Claude Opus 4.8Anthropic28.6
9GPT-5.4OpenAI25.4
10GPT 5.2OpenAI23.0
11Claude Opus 4.7Anthropic19.1
12Claude Opus 4.5Anthropic14.3
13Gemini 3.1 Pro PreviewGoogle14.3
14Claude Opus 4.6Anthropic12.7
15GPT 5OpenAI12.7
16GLM-5.2Zhipu9.5
17Qwen3.7-MaxAlibaba9.5
18Claude Opus 4.1Anthropic7.9
19Gemini 3.5 FlashGoogle4.8
20Claude Sonnet 4.5Anthropic2.4
21Kimi K2.6Kimi2.4

The benchmark's own page