Price one workload
LLM cost calculator
One workload, every listed price. Cheapest first. Full price table.
… models with a listed input and output price
| Rank | Model | Cost | In $/M | Out $/M | Cache rd |
|---|---|---|---|---|---|
| 1 | Cerebras-Llama-4-Scout-17B-16E-Instruct | $0 | free | free | — |
| 2 | Gemma 4 31B IT (free) | $0 | free | free | $0.05 |
| 3 | Kimi K2 | $0 | free | free | $0.40 |
| 4 | MiMo-V2.5-Pro | $0 | free | free | free |
| 5 | Llama 3.2 3b Instruct | $0.036 | $0.0306 | $0.0493 | $0.0153 |
| 6 | gpt-oss-20b | $0.055 | $0.04 | $0.15 | $0.02 |
| 7 | Llama 3.1 8b Instruct | $0.063 | $0.0544 | $0.085 | $0.0272 |
| 8 | Qwen3.5 9B | $0.065 | $0.05 | $0.15 | $0.025 |
| 9 | GLM-5.3-Flash | $0.077 | $0.075 | $0.025 | $0.015 |
| 10 | GPT 5 Nano | $0.09 | $0.05 | $0.40 | $0.005 |
| 11 | Tencent Hy3 | $0.092 | $0.066 | $0.26 | $0.029 |
| 12 | GPT-5.6 Luna | $0.097 | $0.06 | $0.37 | $0.02 |
| 13 | GLM 4.7 Flash Thinking | $0.11 | $0.07 | $0.40 | $0.035 |
| 14 | DeepSeek V4 Flash | $0.118 | $0.098 | $0.196 | $0.028 |
| 15 | Mistral Small 3.2 (Mistral AI) | $0.13 | $0.10 | $0.30 | $0.165 |
| 16 | Qwen3 30B A3B | $0.13 | $0.10 | $0.30 | $0.05 |
| 17 | Qwen3 32B | $0.13 | $0.10 | $0.30 | $0.05 |
| 18 | Stepfun/Step-3.5 Flash | $0.13 | $0.10 | $0.30 | $0.05 |
| 19 | Gemini 2.5 Flash Lite | $0.14 | $0.10 | $0.40 | $0.01 |
| 20 | GPT 4.1 Nano | $0.14 | $0.10 | $0.40 | $0.025 |
| 21 | Gemma 4 26B A4B | $0.158 | $0.12 | $0.38 | $0.06 |
| 22 | Qwen3.8 Flash | $0.182 | $0.14 | $0.42 | $0.016 |
| 23 | Qwen3 235B A22B Thinking 2507 | $0.19 | $0.13 | $0.60 | $0.055 |
| 24 | GLM 4.5 Air | $0.2 | $0.12 | $0.80 | $0.06 |
| 25 | GPT-4o mini | $0.21 | $0.15 | $0.60 | $0.075 |
| 26 | Qwen3 Next 80B A3B Instruct | $0.215 | $0.15 | $0.65 | $0.075 |
| 27 | Qwen3 Next 80B A3B Thinking | $0.215 | $0.15 | $0.65 | $0.075 |
| 28 | Qwen3.8 27B | $0.22 | $0.15 | $0.70 | $0.04 |
| 29 | DeepSeek V3.2 | $0.25 | $0.2174 | $0.326 | $0.06 |
| 30 | X-Ai/Grok 4.1 Fast Non Reasoning | $0.25 | $0.20 | $0.50 | $0.05 |
| 31 | DeepSeek-V3.1 | $0.27 | $0.20 | $0.70 | $0.10 |
| 32 | DeepSeek-V3 | $0.277 | $0.20 | $0.77 | $0.135 |
| 33 | Z-Ai/GLM 4.7 | $0.28 | $0.20 | $0.80 | $0.10 |
| 34 | Step 3.7 Flash | $0.304 | $0.19 | $1.14 | $0.03 |
| 35 | Qwen3.6 Flash | $0.306 | $0.19 | $1.16 | $0.02 |
| 36 | GPT OSS 120B | $0.308 | $0.228 | $0.798 | $2.00 |
| 37 | DeepSeek V3.2 Exp | $0.322 | $0.28 | $0.42 | $0.14 |
| 38 | Minimax/Minimax-M2 | $0.323 | $0.17 | $1.53 | $0.085 |
| 39 | GPT 5.4 Nano | $0.325 | $0.20 | $1.25 | $0.02 |
| 40 | Qwen 3 235B A22B | $0.35 | $0.30 | $0.50 | $0.15 |
| 41 | Qwen3 Coder 480B A35B Instruct | $0.35 | $0.25 | $1.00 | $0.022 |
| 42 | Qwen3 Coder Next | $0.35 | $0.20 | $1.50 | $0.10 |
| 43 | DeepSeek V4.1 Flash | $0.383 | $0.2556 | $1.2778 | $0.0128 |
| 44 | Codestral 2508 | $0.39 | $0.30 | $0.90 | $0.15 |
| 45 | Qwen3-Coder 30B-A3B Instruct | $0.393 | $0.285 | $1.083 | $0.05 |
| 46 | Gemini 3.1 Flash Lite | $0.4 | $0.25 | $1.50 | $0.025 |
| 47 | Qwen3.5 35B A3B | $0.405 | $0.225 | $1.80 | $0.1125 |
| 48 | Gemma 3 27B | $0.41 | $0.342 | $0.684 | $0.1496 |
| 49 | MiniMax M3 Thinking | $0.42 | $0.30 | $1.20 | $0.06 |
| 50 | MiniMax-M2.5 | $0.42 | $0.30 | $1.20 | $0.03 |
| 51 | Qwen3 VL 235B A22B Instruct | $0.42 | $0.30 | $1.20 | $0.15 |
| 52 | Qwen3.6 27B Thinking | $0.427 | $0.203 | $2.24 | $0.1015 |
| 53 | GLM 4.5 | $0.43 | $0.30 | $1.30 | $0.15 |
| 54 | MiniMax M2.7 | $0.441 | $0.315 | $1.26 | $0.1575 |
| 55 | GPT 5 Mini | $0.45 | $0.25 | $2.00 | $0.025 |
| 56 | GPT 5.1 Codex Mini | $0.45 | $0.25 | $2.00 | $0.025 |
| 57 | Qwen3 Coder Flash | $0.45 | $0.30 | $1.50 | $0.15 |
| 58 | Minimax/Minimax-M2.1 | $0.462 | $0.33 | $1.32 | $0.165 |
| 59 | Qwen3.5 27B | $0.486 | $0.27 | $2.16 | $0.135 |
| 60 | Moonshotai/Kimi-K2.5 | $0.49 | $0.30 | $1.90 | $0.15 |
| 61 | Z-AI/GLM 4.6 | $0.49 | $0.35 | $1.40 | $0.175 |
| 62 | Llama 4 Maverick 17B 128E Instruct FP8 | $0.519 | $0.371 | $1.484 | $0.1275 |
| 63 | Qwen 3.6 Plus | $0.52 | $0.325 | $1.95 | $0.0325 |
| 64 | DeepSeek V4 Pro | $0.522 | $0.435 | $0.87 | $0.04 |
| 65 | Qwen3.6 35B-A3B | $0.547 | $0.342 | $2.052 | $0.056 |
| 66 | Gemini 2.5 Flash | $0.55 | $0.30 | $2.50 | $0.03 |
| 67 | Gemini 3.5 Flash Lite | $0.55 | $0.30 | $2.50 | $0.03 |
| 68 | GPT 4.1 Mini | $0.56 | $0.40 | $1.60 | $0.10 |
| 69 | DeepSeek-R1 | $0.57 | $0.40 | $1.70 | $0.20 |
| 70 | DeepSeek V4 Flash Vision Exp | $0.572 | $0.44 | $1.32 | $0.014 |
| 71 | MiMo-V2.5 | $0.6 | $0.40 | $2.00 | $0.08 |
| 72 | Qwen3.5 Plus Thinking | $0.64 | $0.40 | $2.40 | $0.04 |
| 73 | GPT-3.5 Turbo | $0.65 | $0.50 | $1.50 | free |
| 74 | Qwen3.5 397B-A17B | $0.665 | $0.40 | $2.65 | $0.20 |
| 75 | GLM 4.5V Thinking | $0.78 | $0.60 | $1.80 | $0.30 |
| 76 | Qwen3.5 122B A10B Thinking | $0.787 | $0.437 | $3.496 | $0.1038 |
| 77 | Gemini 3 Flash (Preview) | $0.8 | $0.50 | $3.00 | $0.05 |
| 78 | MiniMax-M2.7-highspeed | $0.84 | $0.60 | $2.40 | $0.06 |
| 79 | Kimi K2 Thinking | $0.85 | $0.60 | $2.50 | $0.15 |
| 80 | Qwen3.7 Plus | $1.075 | $0.768 | $3.072 | $0.08 |
Common questions
- How do I estimate the cost of an LLM workload?
- Set input tokens, output tokens and optional cache-read tokens. The table multiplies those by each model's listed rates. Cache-read uses the cache price when a host publishes one. It uses the input price if not.
- What token mix should I use?
- The default is 1 million input tokens and 100 thousand output tokens. That matches a long document plus a short answer. Change the fields. The URL keeps the mix so you can share it.