modelbenchmark.io

What a cache hit saves

Prompt caching prices

What a cache hit saves, and how many reuses pay for the write. Every figure is derived from the prices a host lists. Nothing here is measured. All prices.

Price

1,153 models

DeepSeek V4 FlashDeepSeek$0.098free$0.02871.4%0
GLM-5.2Alibaba$1.40free$0.2681.4%0
DeepSeek V4 ProDeepSeek$0.435free$0.0490.8%0
Kimi K2.6Kimi$0.95free$0.1683.2%0
Kimi K3Alibaba$3.00$3.00$0.3090%0
Kimi K2.7 CodeKimi$0.9006free$0.188179.1%0
GPT OSS 120BOpenAI$0.228$0.18$2.00-777.2%never
GLM-5.1Zhipu$1.40free$0.2681.4%0
GLM-5.3-FlashZhipu$0.075free$0.01580%0
Moonshotai/Kimi-K2.5Alibaba$0.30free$0.1550%0
GLM-5.3Zhipu$1.40free$0.2681.4%0
Claude Haiku 4.5 ThinkingAnthropic$1.00$1.25+25%$0.1090%1
Claude Sonnet 4.6Anthropic$3.00$3.75+25%$0.3090%1
MiniMax-M2.5MiniMax$0.30free$0.0390%0
Claude Opus 4.8Anthropic$5.00$6.25+25%$0.5090%1
GPT 5.5OpenAI$5.00$5.00$0.5090%0
MiniMax M3 ThinkingMiniMax$0.30free$0.0680%0
MiniMax M2.7MiniMax$0.315$0.375+19%$0.157550%1
GLM-5Alibaba$0.89free$0.222675.0%0
Claude 4.7 OpusAnthropic$5.00$6.25+25%$0.5090%1
GPT 5.4OpenAI$2.50$2.50$0.2590%0
Claude Sonnet 4.5Anthropic$3.00$3.75+25%$0.3090%1
gpt-oss-20bOpenAI$0.04$0.07+75%$0.0250%2
DeepSeek-R1DeepSeek$0.40$0.825+106.2%$0.2050%3
Claude 4.6 Opus ThinkingAnthropic$5.00$6.25+25%$0.5090%1
Claude Fable 5Anthropic$10.00$12.50+25%$1.0090%1
GPT-5.6 SolOpenAI$2.00$2.50+25%$0.2090%1
Claude 4.5 OpusAnthropic$5.00$6.25+25%$0.5090%1
Claude Sonnet 5 ThinkingAnthropic$2.00$2.50+25%$0.2090%1
Gemini 2.5 FlashGoogle$0.30$0.30$0.0390%0
GPT-5.6 LunaOpenAI$0.06$0.25+316.7%$0.0266.7%5
Gemini 3.1 Pro (Preview)Google$2.00$0.375$0.2090%0
Gemini 3.5 FlashGoogle$1.50$0.0833$0.1590%0
Gemma 4 31B IT (free)Googlefree$0.50$0.05never
GPT 5.6 TerraOpenAI$2.00$2.50+25%$0.2090%1
Gemini 2.5 ProGoogle$1.25$0.375$0.12590%0
GPT 5.4 MiniOpenAI$0.75$0.75$0.07590%0
Qwen3.7 MaxAlibaba$3.00$3.125+4.2%$0.5083.3%1
Qwen3.8 Max ThinkingAlibaba$2.00$2.50+25%$0.2587.5%1
Claude Opus 5Anthropic$5.00$6.25+25%$0.5090%1
Z-Ai/GLM 4.7Zhipu$0.20free$0.1050%0
GPT 5 MiniOpenAI$0.25$0.25$0.02590%0
Qwen3.7 PlusAlibaba$0.768$0.50$0.0889.6%0
OpenAI/GPT-5OpenAI$1.25$1.25$0.12590%0
Gemini 3 Flash (Preview)Google$0.50$0.0833$0.0590%0
GPT 5.4 NanoOpenAI$0.20$0.20$0.0290%0
Grok 4.6xAI$2.00$2.00$0.5075%0
Llama-3.3-70B-InstructMeta$1.254$0.90$0.02598%0
OpenAI/GPT-5.2OpenAI$1.75$1.75$0.17590%0
Qwen3.5 397B-A17BAlibaba$0.40$0.75+87.5%$0.2050%2
DeepSeek V4.1 FlashDeepSeek$0.2556$0.27+5.6%$0.012895.0%1
GPT-5.1 (2025-11-13)OpenAI$1.25$1.25$0.12590%0
Grok 4.3xAI$1.25$1.25$0.2084%0
Z-AI/GLM 4.6Zhipu$0.35free$0.17550%0
Gemini 3.1 Flash LiteGoogle$0.25$0.0833$0.02590%0
GPT 5.3 CodexOpenAI$1.75$1.75$0.17590%0
Gemini 3.6 FlashGoogle$0.75$0.0417$0.07590%0
Gemini 3.7 FlashGoogle$0.75$0.0417$0.07590%0
GPT 4.1OpenAI$2.00$2.00$0.5075%0
GPT 5 NanoOpenAI$0.05$0.05$0.00590%0
Claude 4.1 OpusAnthropic$15.00$18.75+25%$1.5090%1
Gemini 3.5 Flash LiteGoogle$0.30$0.0833$0.0390%0
GPT 4.1 MiniOpenAI$0.40$0.40$0.1075%0
Grok 4.5xAI$2.00$2.00$0.5075%0
Kimi K2 ThinkingKimi$0.60$1.10+83.3%$0.1575%2
Qwen 3.6 PlusAlibaba$0.325$0.4062+25.0%$0.032590.0%1
Qwen3 Max PreviewAlibaba$1.2002$1.50+25.0%$0.600150%1
Qwen3 Next 80B A3B InstructAlibaba$0.15$0.20+33.3%$0.07550%1
Qwen3.8 27BAlibaba$0.15$0.625+316.7%$0.0473.3%5
Claude Fable 5.1Anthropic$10.00$12.50+25%$0.2597.5%1
DeepSeek V3.2 ExpDeepSeek$0.28$0.1450%
MiMo-V2.5-Profreefreefreenever
Minimax/Minimax-M2.1MiniMax$0.33free$0.16550%0
Qwen3 235B A22B Instruct 2507Alibaba$1.026$1.20+17.0%$0.06593.7%1
Qwen3.6 35B-A3BAlibaba$0.342$0.175$0.05683.6%0
DeepSeek V3.2DeepSeek$0.2174$0.375+72.5%$0.0672.4%2
Gemini 2.5 Flash LiteGoogle$0.10$0.10$0.0190%0
GPT-4o (2024-08-06)OpenAI$2.50$2.50$1.2550%0
Qwen3.6 27B ThinkingAlibaba$0.203$0.60+195.6%$0.101550%4
GPT-4o miniOpenAI$0.15$0.15$0.07550%0
Qwen3 32BAlibaba$0.10$0.90+800%$0.0550%16
GLM 4.5Zhipu$0.30free$0.1550%0
GPT 4.1 NanoOpenAI$0.10$0.10$0.02575%0
MiMo-V2.5$0.40free$0.0880%0
Gemini 3.8 FlashGoogle$0.75$0.0417$0.07590%0
Gemma 4 26B A4BGoogle$0.12$0.0650%
GPT 5.2 CodexOpenAI$1.75$1.75$0.17590%0
GPT 6 AstraOpenAI$10.00$12.50+25%$1.0090%1
Inkling ThinkingDatabricks$1.00$1.00$0.1783%0
Qwen3 Next 80B A3B ThinkingAlibaba$0.15$0.07550%
Qwen3.5 9BAlibaba$0.05$0.30+500%$0.02550%10
Qwen3.8 FlashAlibaba$0.14$0.20+42.9%$0.01688.6%1
Claude 4 Sonnet ThinkingAnthropic$3.00$3.75+25%$0.3090%1
GLM 4.5 AirZhipu$0.12free$0.0650%0
GLM 4.7 Flash ThinkingZhipu$0.07free$0.03550%0
OpenAI o4-miniOpenAI$1.10free$0.5550%0
Qwen3 235B A22B Thinking 2507Alibaba$0.13$0.22+69.2%$0.05557.7%2
Qwen3 VL 235B A22B InstructAlibaba$0.30$0.2625$0.1550%0
Qwen3-Coder 30B-A3B InstructAlibaba$0.285$0.075$0.0582.5%0
Qwen3.6 FlashAlibaba$0.19$0.24+26.3%$0.0289.5%1
DeepSeek-V3DeepSeek$0.20$0.27+35%$0.13532.5%2
GPT 5.1 Codex MiniOpenAI$0.25$0.25$0.02590%0
Minimax/Minimax-M2MiniMax$0.17$0.375+120.6%$0.08550%3
Mistral Large 2411Mistral$2.006$0.2090%
OpenAI o3OpenAI$2.00free$1.0050%0
Qwen3.5 Plus ThinkingAlibaba$0.40free$0.0490%0
Tencent Hy3$0.066$0.16+142.4%$0.02956.1%3
GPT 5.1 CodexOpenAI$1.25$0.12590%
GPT-5.4 ProOpenAI$30.00$3.0090%
GPT-5.5 ProOpenAI$30.00$30.00$3.0090%0
Mistral Small 3.2 (Mistral AI)Mistral$0.10$0.165-65%
OpenAI o3-miniOpenAI$1.10free$0.5550%0
Qwen3 Coder NextAlibaba$0.20free$0.1050%0
Qwen3.5 122B A10B ThinkingAlibaba$0.437$0.625+43%$0.103876.2%1
Qwen3.5 35B A3BAlibaba$0.225$0.112550%
DeepSeek V4 Flash Vision ExpDeepSeek$0.44$0.27$0.01496.8%0
GLM 5 TurboZhipu$1.20free$0.2480%0
Grok Build 0.1xAI$1.00$1.00$0.2080%0
Qwen3 Coder 480B A35B InstructAlibaba$0.25$0.02291.2%
Qwen3 Coder PlusAlibaba$1.00$1.25+25%$0.5050%1
Step 3.7 Flash$0.19free$0.0384.2%0
Claude 4 Opus ThinkingAnthropic$15.00$18.75+25%$1.5090%1
Gemma 3 27BGoogle$0.342$0.125$0.149656.3%0
GLM 5V Turbo ThinkingZhipu$1.20free$0.2480%0
MiniMax-M2.7-highspeedMiniMax$0.60$0.375$0.0690%0
Qwen3 VL 235B A22B ThinkingAlibaba$0.50$0.2550%
Claude 3.7 SonnetAnthropic$3.00$3.75+25%$0.3090%1
Codestral 2508Mistral$0.30$0.1550%
Gemini-3-ProGoogle$1.60$2.50+56.2%$0.1690%1
Kimi K2Kimifree$0.40
Llama 3.1 8b InstructMeta$0.0544$0.20+267.6%$0.027250%6
Qwen3.5 27BAlibaba$0.27$0.13550%
X-Ai/Grok 4.1 Fast Non ReasoningxAI$0.20$0.0575%
GPT 5 ProOpenAI$15.00$1.5090%
GPT 5.1 Codex MaxOpenAI$2.50$1.25$0.2590%0
GPT-3.5 TurboOpenAI$0.50free100%
Llama 3.2 3b InstructMeta$0.0306$0.10+226.8%$0.015350%5
OpenAI o1OpenAI$15.00free$7.5050%0
Qwen 3 235B A22BAlibaba$0.30$0.09$0.1550%0
Qwen3 30B A3BAlibaba$0.10$0.0550%
Stepfun/Step-3.5 Flash$0.10$0.0550%
DeepSeek-V3.1DeepSeek$0.20$0.56+180%$0.1050%4
Gemini 3 Pro ImageGoogle$2.00$4.50+125%$0.2090%2
GLM 4.5V ThinkingZhipu$0.60free$0.3050%0
GPT-4 TurboOpenAI$10.00
Llama 4 Maverick 17B 128E Instruct FP8Meta$0.371$0.30$0.127565.6%0
Qwen3 Coder FlashAlibaba$0.30$0.375+25%$0.1550%1
DeepSeek/DeepSeek-V3.1-TerminusDeepSeek$0.25$0.12550%
Gemma 3 12B ITGoogle$0.272$0.13650%
Nano Banana 2Google$0.50$0.0590%

149 of these models list a cache-read price and 128 list a cache-write price. Break-even is blank for the rest: without price_cache_write the write premium is unknown, and assuming it is zero would put a confident 0 in 21 rows nobody measured. There is no TTL column and no minimum-token column because no source we ingest publishes either field.

A worked example

One million prefix tokens on Claude Haiku 4.5 Thinking, sent once and reused 10 times. Cached costs $1.25 to write plus 10 reads at $0.10, so $2.25. Uncached costs 11 sends at $1.00, so $11.00. Break-even for this model is 1 later request.

Common questions

What is prompt caching?
A host stores the tokens at the start of your prompt. A later request that repeats that prefix is charged a cache-read rate instead of the input rate. The saving is on input tokens only. Output is charged as usual.
How much does a cache hit save?
The discount column is (input price − cache-read price) ÷ input price, from the prices each host lists. It is not a measurement. Most hosts list a cache read at one tenth of their input price.
What is the break-even column?
The smallest number of later requests at which caching a prefix costs less than sending it uncached every time. Writing once and reading n times costs (cache write + n × cache read). Not caching costs (1 + n) × input. The column is blank when the host lists no cache-write price, because the write premium is then unknown.
Why is there no cache TTL or minimum-token column?
No source this site ingests publishes either one. Checked on 2026-09-02 against every corpus row: the fields available are the cache-read price, the cache-write price, a longer-TTL write price at a few hosts, and a caching capability flag. A TTL column would be invented, so there is none.
Do all hosts of one model charge the same cache price?
No. A price belongs to a model and a host. This table shows one listing per model. Open the model page to see every host.