modelbenchmark.io

How the score is built

Benchmarks

A benchmark earns its place in the composite by discriminating. We measure two things for each one: how far apart the models are on it, and how far apart two measurements of the same model are. The weight is the share of the first that survives the second.

Weight = (spread² − noise²) / spread²A benchmark every model passes has almost no spread left once its own measurement error is subtracted, so it stops counting on its own. No weight on this page was chosen by hand.

Sources

SourceWho ran itWhere
Epoch AIindependent runsepoch.ai/benchmarks
SWE-benchbenchmark authorswww.swebench.com
LiveBenchbenchmark authorslivebench.ai
Aiderbenchmark authorsaider.chat/docs/leaderboards

Epoch AI data is used under CC-BY.

In the composite

Measured, but not counted