modelbenchmark.io

Best AI for tool calling

Models whose listings report tool calling. Ranked by composite when scored, else by output price. Full leaderboard.

RankModelLabCompositeContextOut $/M
1GPT 6 AstraOpenAI100th1.1M$50.00
2Claude Fable 5Anthropic99th1M$50.00
3Claude Fable 5.1Anthropic99th1M$50.00
4GPT-5.5 ProOpenAI98th1.1M$180.00
5Claude Opus 5Anthropic96th1M$25.00
6GPT-5.4 ProOpenAI95th1.1M$180.00
7Gemini 3.8 FlashGoogle95th1M$3.75
8GPT-5.6 SolOpenAI94th1.1M$10.00
9DeepSeek V4.1 FlashDeepSeek94th1M$1.2778
10Gemini 3.7 FlashGoogle94th1M$3.75
11GPT 5.5OpenAI93rd1.1M$30.00
12Gemini-3-ProGoogle91st1M$9.60
13MiniMax-M2.5MiniMax91st205K$1.20
14GPT 5.6 TerraOpenAI90th1.1M$12.00
15Qwen3.8 Max ThinkingAlibaba89th991K$6.00
16Claude Opus 4.8Anthropic89th1M$25.00
17GPT 5.3 CodexOpenAI89th400K$14.00
18GPT 5.4OpenAI88th1.1M$15.00
19Grok 4.6xAI87th500K$6.00
20Gemini 3.1 Pro (Preview)Google86th1M$12.00
21Grok 4.5xAI85th500K$6.00
22Gemini 3 Flash (Preview)Google84th1M$3.00
23Gemini 3.6 FlashGoogle84th1M$3.75
24Gemini 3.5 FlashGoogle84th1M$9.00
25Qwen3.7 PlusAlibaba84th1M$3.072
26Claude 4.7 OpusAnthropic83rd1M$25.00
27GLM-5.3Zhipu82nd1M$4.40
28Claude Sonnet 5 ThinkingAnthropic80th1M$10.00
29GPT-5.6 LunaOpenAI80th1.1M$0.37
30DeepSeek V4 Flash Vision ExpDeepSeek80th1M$1.32
31OpenAI/GPT-5.2OpenAI79th400K$14.00
32Qwen3.7 MaxAlibaba79th1M$9.00
33GPT 5 ProOpenAI78th400K$120.00
34DeepSeek-R1DeepSeek78th128K$1.70
35OpenAI/GPT-5OpenAI78th400K$10.00
36GLM-5.1Zhipu77th200K$4.40
37GPT 5.1 CodexOpenAI77th400K$10.00
38OpenAI o3OpenAI76th200K$8.00
39GPT-5.1 (2025-11-13)OpenAI75th400K$10.00
40Claude 4.6 Opus ThinkingAnthropic74th1M$25.00
41GPT 5.2 CodexOpenAI73rd400K$14.00
42Qwen3.6 35B-A3BAlibaba73rd262K$2.052
43Kimi K2Kimi72nd128Kfree
44Kimi K2.6Kimi72nd262K$4.00
45Moonshotai/Kimi-K2.5Alibaba71st256K$1.90
46Qwen3.5 397B-A17BAlibaba70th262K$2.65
47DeepSeek R1 Distill Llama 70BDeepSeek69th33K$0.99
48DeepSeek V3.2DeepSeek68th128K$0.326
49Minimax/Minimax-M2MiniMax68th200K$1.53
50GLM-5Alibaba66th205K$3.2647
51GLM-5.2Alibaba66th1M$4.40
52Qwen3.5 Plus ThinkingAlibaba65th984K$2.40
53Kimi K2.7 CodeKimi63rd262K$4.389
54Qwen3.5 35B A3BAlibaba63rd260K$1.80
55GLM-5.3-FlashZhipu61st1M$0.025
56Inkling ThinkingDatabricks61st1M$4.05
57OpenAI o1OpenAI59th200K$60.00
58Qwen3 Max PreviewAlibaba58th256K$6.001
59Gemini 2.5 ProGoogle57th1M$10.00
60Gemini 3.1 Flash LiteGoogle56th1M$1.50
61Claude 4.5 OpusAnthropic55th200K$25.00
62OpenAI o4-miniOpenAI55th200K$4.40
63GPT 5 MiniOpenAI54th400K$2.00
64Claude Sonnet 4.5Anthropic54th200K$15.00
65Z-Ai/GLM 4.7Zhipu51st200K$0.80
66GPT 5.4 MiniOpenAI50th400K$4.50
67Qwen3.5 9BAlibaba48th256K$0.15
68GPT OSS 120BOpenAI47th131K$0.798
69GPT 5.4 NanoOpenAI46th400K$1.25
70Qwen3.6 FlashAlibaba46th992K$1.16
71Qwen3.5 122B A10B ThinkingAlibaba45th131K$3.496
72Claude 4.1 OpusAnthropic43rd200K$75.00
73GLM 4.7 Flash ThinkingZhipu43rd200K$0.40
74OpenAI o3-miniOpenAI42nd200K$4.40
75GPT 4.1OpenAI41st1M$8.00
76Gemini 2.5 FlashGoogle40th1M$2.50
77GPT 5 NanoOpenAI40th400K$0.40
78Gemini 3.5 Flash LiteGoogle39th1M$2.50
79Z-AI/GLM 4.6Zhipu38th200K$1.40
80DeepSeek-V3DeepSeek35th128K$0.77
81DeepSeek V4 ProDeepSeek35th1M$0.87
82GLM 4.5Zhipu34th131K$1.30
83GPT 4.1 MiniOpenAI31st1M$1.60
84GPT 4.1 NanoOpenAI22nd1M$0.40
85GPT-4o (2024-08-06)OpenAI21st128K$10.00
86Mistral Large 2411Mistral15th128K$6.001
87GPT-4o miniOpenAI15th128K$0.60
88Grok Build 0.1xAI12th256K$2.00
89GPT-3.5 TurboOpenAI9th16K$1.50
90Grok 4.3xAI0th1M$2.50
91Cerebras-Llama-4-Scout-17B-16E-InstructMeta128Kfree
92Gemma 4 31B IT (free)Google262Kfree
93MiMo-V2.5-Pro1Mfree
94Llama 3.2 3b InstructMeta131K$0.0493
95Llama 3.1 8b InstructMeta131K$0.085
96gpt-oss-20bOpenAI128K$0.15
97DeepSeek V4 FlashDeepSeek1M$0.196
98Tencent Hy3262K$0.26
99Mistral Small 3.2 (Mistral AI)Mistral128K$0.30
100Qwen3 30B A3BAlibaba40K$0.30
101Qwen3 32BAlibaba40K$0.30
102Stepfun/Step-3.5 Flash64K$0.30
103Gemma 4 26B A4BGoogle262K$0.38
104Gemini 2.5 Flash LiteGoogle1M$0.40
105DeepSeek V3.2 ExpDeepSeek164K$0.42
106Qwen3.8 FlashAlibaba992K$0.42
107Qwen 3 235B A22BAlibaba128K$0.50
108X-Ai/Grok 4.1 Fast Non ReasoningxAI2M$0.50
109Qwen3 235B A22B Thinking 2507Alibaba262K$0.60
110Qwen3 Next 80B A3B InstructAlibaba131K$0.65
111Qwen3 Next 80B A3B ThinkingAlibaba131K$0.65
112Gemma 3 27BGoogle40K$0.684
113DeepSeek-V3.1DeepSeek128K$0.70
114Qwen3.8 27BAlibaba262K$0.70
115GLM 4.5 AirZhipu131K$0.80
116Codestral 2508Mistral256K$0.90
117Qwen3 Coder 480B A35B InstructAlibaba262K$1.00
118Qwen3-Coder 30B-A3B InstructAlibaba128K$1.083
119Step 3.7 Flash262K$1.14
120MiniMax M3 ThinkingMiniMax512K$1.20
121Qwen3 VL 235B A22B InstructAlibaba131K$1.20
122Llama-3.3-70B-InstructMeta100K$1.254
123MiniMax M2.7MiniMax205K$1.26
124Minimax/Minimax-M2.1MiniMax205K$1.32
125Llama 4 Maverick 17B 128E Instruct FP8Meta131K$1.484
126Qwen3 Coder FlashAlibaba128K$1.50
127Qwen3 Coder NextAlibaba262K$1.50
128GLM 4.5V ThinkingZhipu66K$1.80
129Qwen 3.6 PlusAlibaba992K$1.95
130GPT 5.1 Codex MiniOpenAI400K$2.00
131MiMo-V2.51M$2.00
132Qwen3.5 27BAlibaba260K$2.16
133Qwen3.6 27B ThinkingAlibaba260K$2.24
134MiniMax-M2.7-highspeedMiniMax205K$2.40
135Kimi K2 ThinkingKimi256K$2.50
136Qwen3 235B A22B Instruct 2507Alibaba262K$3.078
137GLM 5 TurboZhipu203K$4.00
138GLM 5V Turbo ThinkingZhipu203K$4.00
139Claude Haiku 4.5 ThinkingAnthropic200K$5.00
140Qwen3 Coder PlusAlibaba128K$5.00
141Qwen3 VL 235B A22B ThinkingAlibaba131K$6.00
142Gemini 3 Pro ImageGoogle66K$12.00
143Claude 3.7 SonnetAnthropic200K$15.00
144Claude 4 Sonnet ThinkingAnthropic1M$15.00
145Claude Sonnet 4.6Anthropic1M$15.00
146Kimi K3Alibaba1M$15.00
147GPT 5.1 Codex MaxOpenAI400K$20.00
148GPT-4 TurboOpenAI128K$30.00
149Claude 4 Opus ThinkingAnthropic200K$75.00

Common questions

What is the best ai for tool calling right now?
Models whose listings report tool calling. Ranked by composite when scored, else by output price.
How is this ranking computed?
The set is a spec filter on the catalog. Scored models rank by the overall composite. Unscored models follow by the spec itself.