modelbenchmark.io

Mercury 2.5 Preview

mercury-2-5

Compare

Mercury 2.5 Preview is Inception's latest and most intelligent diffusion language model. Instead of generating tokens strictly one at a time, it produces and refines multiple tokens in parallel, reaching up to 1,107 tokens per second on standard GPUs. It delivers a 10+ point intelligence gain over Mercury 2, with tunable reasoning, parallel tool calls, schema-aligned JSON output, and a 260K context window. It is built for latency-sensitive production work such as search agents, voice pipelines, customer support, rapid coding iteration, and coding subagents.

Specification

most-agreed values

Context
260K
Max output
66K
Released
2026-09-01
Knowledge cutoff
2025-11-01
Retires
Open weights
no
Input
text
Output
text

Price

US dollars per million tokens · most-agreed

Input
$0.04
Output
$0.15
Cache read
$0.004

Available from 7 hosts

HostIn $/MOut $/MCache rdCache wrContextOutputRetires
inception$0.04$0.15$0.004260K66K
nano-gpt · inception$0.04$0.15$0.004260K66K
openrouter · inception$0.04$0.15$0.004260K66K
vercel · inception$0.04$0.15$0.004260K66K
venice$0.05$0.1875$0.005260K66K
inception · inception$0.20$0.75260K66K
kilo · inception$0.20$0.75$0.02260K66K

Price history

append-only observations · a listing writes a row only when its price moves

DateHostField$/M
2026-09-09veniceCache read$0.005
2026-09-10veniceCache read$0.025
2026-09-11veniceCache read$0.005
2026-09-09veniceInput$0.05
2026-09-10veniceInput$0.25
2026-09-11veniceInput$0.05
2026-09-09veniceOutput$0.1875
2026-09-10veniceOutput$0.9375
2026-09-11veniceOutput$0.1875