modelbenchmark.io

Ranked by composite score

AI model leaderboard

One number per model, built from 16 published benchmarks across 4 sources. Each benchmark is weighted by how much of its spread survives its own measurement error, so a benchmark that no longer separates models stops counting. How the weights are computed. The quality rank uses scores only. Best per dollar is quality over output price. Speed is not in any rank. We do not sample live APIs.

Scope
Rank by
Weights
Price
Evidence
View

79 models ranked · 3 or more benchmarks, from 2 or more sources · percentile is out of all 334 scored models either way

RankModelLabCompositeEvidenceEffort rangeOut $/M
1GPT-6 AstraOpenAI100th11 bench · 2 src9.8 pts$50.00
2Claude Fable 5Anthropic99th14 bench · 2 src7.1 pts$50.00
3Claude Fable 5.1Anthropic99th10 bench · 2 src$50.00
4Claude Opus 5Anthropic96th9 bench · 2 src14.0 pts$25.00
5Gemini 3.8 FlashGoogle95th8 bench · 2 src$3.75
6GPT-5.6 SolOpenAI94th12 bench · 2 src16.2 pts$10.00
7Gemini 3.7 FlashGoogle94th8 bench · 2 src$3.75
8GPT-5.5OpenAI93rd16 bench · 2 src21.8 pts$30.00
9DeepSeek V4 Pro 0813DeepSeek91st8 bench · 2 src$4.0077
10Gemini 3 Pro PreviewGoogle91st6 bench · 2 src6.2 pts$9.60
11GPT-5.6 TerraOpenAI90th8 bench · 2 src17.9 pts$12.00
12Qwen3.8 MaxAlibaba89th8 bench · 2 src$6.00
13Claude Opus 4.8Anthropic89th13 bench · 2 src3.8 pts$25.00
14GPT-5.4OpenAI88th13 bench · 2 src16.7 pts$15.00
15Grok 4.6xAI87th9 bench · 2 src1.1 pts$6.00
16Gemini 3.1 Pro PreviewGoogle86th12 bench · 2 src2.7 pts$12.00
17Grok 4.5xAI85th7 bench · 2 src$6.00
18Kimi K3Kimi85th8 bench · 2 src16.5 ptsfree
19Gemini 3 Flash PreviewGoogle84th10 bench · 2 src2.8 pts$3.00
20Gemini 3.6 FlashGoogle84th8 bench · 2 src8.1 pts$3.75
21Gemini 3.5 FlashGoogle84th14 bench · 3 src7.0 pts$9.00
22Claude Opus 4.7Anthropic83rd13 bench · 2 src11.1 pts$25.00
23GLM-5.3Zhipu82nd8 bench · 2 src$4.40
24Grok 4xAI82nd6 bench · 2 src$15.00
25DeepSeek V4 Flash 0731DeepSeek81st8 bench · 2 src$0.50
26Claude Sonnet 5Anthropic80th8 bench · 2 src12.5 pts$10.00
27GPT-5.6 LunaOpenAI80th8 bench · 2 src19.0 pts$0.37
28GPT 5.2OpenAI79th12 bench · 3 src14.3 pts$14.00
29Qwen3.7-MaxAlibaba79th10 bench · 2 src$9.00
30DeepSeek R1DeepSeek78th4 bench · 2 src$1.70
31GPT 5OpenAI78th13 bench · 3 src8.8 pts$10.00
32o3OpenAI76th11 bench · 3 src4.2 pts$8.00
33GPT 5.1OpenAI75th8 bench · 2 src8.3 pts$10.00
34Claude Opus 4.6Anthropic74th11 bench · 2 src4.3 pts$25.00
35Kimi K2 InstructKimi72nd7 bench · 3 srcfree
36Kimi K2.6Kimi72nd12 bench · 2 src$4.00
37Kimi K2.5Kimi71st7 bench · 2 src3.0 pts$1.90
38Claude Opus 4Anthropic70th7 bench · 2 src4.6 pts$75.00
39DeepSeek-V3.2-ExpDeepSeek68th7 bench · 3 src10.0 pts$0.326
40Qwen3.6 PlusAlibaba67th9 bench · 2 src$3.00
41Claude Sonnet 4Anthropic66th7 bench · 3 src11.6 pts$15.00
42GLM 5Zhipu66th6 bench · 2 src0.7 pts$3.2647
43GLM-5.2Zhipu66th10 bench · 2 src4.5 pts$4.40
44Claude Sonnet 4.6Anthropic64th8 bench · 2 src9.3 pts$15.94
45Claude 3.7 SonnetAnthropic64th7 bench · 3 src12.5 pts$15.00
46Kimi K2.7 CodeKimi63rd7 bench · 2 src$4.389
47GPT-4.5OpenAI62nd4 bench · 2 src$150.00
48Qwen3.6 27BAlibaba62nd5 bench · 2 srcfree
49GLM-5.3-FlashZhipu61st7 bench · 2 src$0.025
50InklingThinking Machines61st7 bench · 2 src$4.05
51Grok 3 Betanot in catalogxAI59th6 bench · 2 src
52o1-2024-12-17OpenAI59th9 bench · 3 src$60.00
53Gemini 2.5 ProGoogle57th8 bench · 2 src4.0 pts$10.00
54Claude Opus 4.5Anthropic55th11 bench · 2 src3.3 pts$25.00
55o4-miniOpenAI55th12 bench · 3 src$4.40
56GPT 5 miniOpenAI54th11 bench · 2 src3.3 pts$2.00
57MiniMax M3MiniMax53rd5 bench · 2 src$2.40
58GPT-5.4 miniOpenAI50th10 bench · 2 src3.3 pts$4.50
59o1-miniOpenAI49th5 bench · 2 src$4.40
60gpt-oss-120bOpenAI47th6 bench · 3 src$0.798
61GPT-5.4 nanoOpenAI46th10 bench · 2 src13.0 pts$1.25
62O3 MiniOpenAI42nd12 bench · 3 src6.6 pts$4.40
63GPT 4.1OpenAI41st10 bench · 3 src9.0 pts$8.00
64Qwen3 32BAlibaba41st4 bench · 2 src$0.55
65QwQ 32BAlibaba41st4 bench · 2 src$0.861
66GPT 5 nanoOpenAI40th11 bench · 2 src10.8 pts$0.40
67Gemini 3.5 Flash-LiteGoogle39th7 bench · 2 src8.0 pts$2.50
68GLM 4.6Zhipu38th3 bench · 2 src$1.40
69Gemini 2.0 Flash ThinkingGoogle37th6 bench · 3 src$0.42
70GLM 4.5Zhipu34th3 bench · 2 src$1.30
71GPT 4.1 miniOpenAI31st10 bench · 3 src$1.60
72Llama 4 Maverick InstructMeta29th6 bench · 3 src$0.60
73Gemma 3 27BGoogle28th5 bench · 2 srcfree
74GPT-4.1 nanoOpenAI22nd6 bench · 2 src$0.40
75GPT 4oOpenAI21st9 bench · 3 src2.0 pts$10.00
76Llama 4 Scout InstructMeta20th5 bench · 2 src$0.46
77GPT-4o miniOpenAI15th8 bench · 2 src$0.60
78Claude 3 OpusAnthropic14th6 bench · 2 src$75.00
79Claude 2not in catalogAnthropic2nd4 bench · 2 src

Effort range is the median spread between the same model's reasoning-effort settings, in points. It is larger than the disagreement between sources, so a score quoted without its configuration is under-specified.

Evidence counts the benchmarks behind the number and the independent sources behind those. The percentile denominator never changes with the filter.

Open weights uses the catalog flag. Closed is every catalog model without that flag. A scored model with no catalog row is neither.

Paid is a listed output price above zero. Free is a listed output price of zero. An open-weight model can have a paid API.