modelbenchmark.io

Price one workload

LLM cost calculator

One workload, every listed price. Cheapest first. Full price table.

models with a listed input and output price

RankModelCostIn $/MOut $/MCache rd
1Cerebras-Llama-4-Scout-17B-16E-Instruct$0freefree
2Gemma 4 31B IT (free)$0freefree$0.05
3Kimi K2$0freefree$0.40
4MiMo-V2.5-Pro$0freefreefree
5Llama 3.2 3b Instruct$0.036$0.0306$0.0493$0.0153
6gpt-oss-20b$0.055$0.04$0.15$0.02
7Llama 3.1 8b Instruct$0.063$0.0544$0.085$0.0272
8Qwen3.5 9B$0.065$0.05$0.15$0.025
9GLM-5.3-Flash$0.077$0.075$0.025$0.015
10GPT 5 Nano$0.09$0.05$0.40$0.005
11Tencent Hy3$0.092$0.066$0.26$0.029
12GPT-5.6 Luna$0.097$0.06$0.37$0.02
13GLM 4.7 Flash Thinking$0.11$0.07$0.40$0.035
14DeepSeek V4 Flash$0.118$0.098$0.196$0.028
15Mistral Small 3.2 (Mistral AI)$0.13$0.10$0.30$0.165
16Qwen3 30B A3B$0.13$0.10$0.30$0.05
17Qwen3 32B$0.13$0.10$0.30$0.05
18Stepfun/Step-3.5 Flash$0.13$0.10$0.30$0.05
19Gemini 2.5 Flash Lite$0.14$0.10$0.40$0.01
20GPT 4.1 Nano$0.14$0.10$0.40$0.025
21Gemma 4 26B A4B$0.158$0.12$0.38$0.06
22Qwen3.8 Flash$0.182$0.14$0.42$0.016
23Qwen3 235B A22B Thinking 2507$0.19$0.13$0.60$0.055
24GLM 4.5 Air$0.2$0.12$0.80$0.06
25GPT-4o mini$0.21$0.15$0.60$0.075
26Qwen3 Next 80B A3B Instruct$0.215$0.15$0.65$0.075
27Qwen3 Next 80B A3B Thinking$0.215$0.15$0.65$0.075
28Qwen3.8 27B$0.22$0.15$0.70$0.04
29DeepSeek V3.2$0.25$0.2174$0.326$0.06
30X-Ai/Grok 4.1 Fast Non Reasoning$0.25$0.20$0.50$0.05
31DeepSeek-V3.1$0.27$0.20$0.70$0.10
32DeepSeek-V3$0.277$0.20$0.77$0.135
33Z-Ai/GLM 4.7$0.28$0.20$0.80$0.10
34Step 3.7 Flash$0.304$0.19$1.14$0.03
35Qwen3.6 Flash$0.306$0.19$1.16$0.02
36GPT OSS 120B$0.308$0.228$0.798$2.00
37DeepSeek V3.2 Exp$0.322$0.28$0.42$0.14
38Minimax/Minimax-M2$0.323$0.17$1.53$0.085
39GPT 5.4 Nano$0.325$0.20$1.25$0.02
40Qwen 3 235B A22B$0.35$0.30$0.50$0.15
41Qwen3 Coder 480B A35B Instruct$0.35$0.25$1.00$0.022
42Qwen3 Coder Next$0.35$0.20$1.50$0.10
43DeepSeek V4.1 Flash$0.383$0.2556$1.2778$0.0128
44Codestral 2508$0.39$0.30$0.90$0.15
45Qwen3-Coder 30B-A3B Instruct$0.393$0.285$1.083$0.05
46Gemini 3.1 Flash Lite$0.4$0.25$1.50$0.025
47Qwen3.5 35B A3B$0.405$0.225$1.80$0.1125
48Gemma 3 27B$0.41$0.342$0.684$0.1496
49MiniMax M3 Thinking$0.42$0.30$1.20$0.06
50MiniMax-M2.5$0.42$0.30$1.20$0.03
51Qwen3 VL 235B A22B Instruct$0.42$0.30$1.20$0.15
52Qwen3.6 27B Thinking$0.427$0.203$2.24$0.1015
53GLM 4.5$0.43$0.30$1.30$0.15
54MiniMax M2.7$0.441$0.315$1.26$0.1575
55GPT 5 Mini$0.45$0.25$2.00$0.025
56GPT 5.1 Codex Mini$0.45$0.25$2.00$0.025
57Qwen3 Coder Flash$0.45$0.30$1.50$0.15
58Minimax/Minimax-M2.1$0.462$0.33$1.32$0.165
59Qwen3.5 27B$0.486$0.27$2.16$0.135
60Moonshotai/Kimi-K2.5$0.49$0.30$1.90$0.15
61Z-AI/GLM 4.6$0.49$0.35$1.40$0.175
62Llama 4 Maverick 17B 128E Instruct FP8$0.519$0.371$1.484$0.1275
63Qwen 3.6 Plus$0.52$0.325$1.95$0.0325
64DeepSeek V4 Pro$0.522$0.435$0.87$0.04
65Qwen3.6 35B-A3B$0.547$0.342$2.052$0.056
66Gemini 2.5 Flash$0.55$0.30$2.50$0.03
67Gemini 3.5 Flash Lite$0.55$0.30$2.50$0.03
68GPT 4.1 Mini$0.56$0.40$1.60$0.10
69DeepSeek-R1$0.57$0.40$1.70$0.20
70DeepSeek V4 Flash Vision Exp$0.572$0.44$1.32$0.014
71MiMo-V2.5$0.6$0.40$2.00$0.08
72Qwen3.5 Plus Thinking$0.64$0.40$2.40$0.04
73GPT-3.5 Turbo$0.65$0.50$1.50free
74Qwen3.5 397B-A17B$0.665$0.40$2.65$0.20
75GLM 4.5V Thinking$0.78$0.60$1.80$0.30
76Qwen3.5 122B A10B Thinking$0.787$0.437$3.496$0.1038
77Gemini 3 Flash (Preview)$0.8$0.50$3.00$0.05
78MiniMax-M2.7-highspeed$0.84$0.60$2.40$0.06
79Kimi K2 Thinking$0.85$0.60$2.50$0.15
80Qwen3.7 Plus$1.075$0.768$3.072$0.08

Common questions

How do I estimate the cost of an LLM workload?
Set input tokens, output tokens and optional cache-read tokens. The table multiplies those by each model's listed rates. Cache-read uses the cache price when a host publishes one. It uses the input price if not.
What token mix should I use?
The default is 1 million input tokens and 100 thousand output tokens. That matches a long document plus a short answer. Change the fields. The URL keeps the mix so you can share it.