Mercury 2.5 Preview
mercury-2-5
Mercury 2.5 Preview is Inception's latest and most intelligent diffusion language model. Instead of generating tokens strictly one at a time, it produces and refines multiple tokens in parallel, reaching up to 1,107 tokens per second on standard GPUs. It delivers a 10+ point intelligence gain over Mercury 2, with tunable reasoning, parallel tool calls, schema-aligned JSON output, and a 260K context window. It is built for latency-sensitive production work such as search agents, voice pipelines, customer support, rapid coding iteration, and coding subagents.
Specification
most-agreed values
Context
260K
Max output
66K
Released
2026-09-01
Knowledge cutoff
2025-11-01
Retires
—
Open weights
no
Input
text
Output
text
Price
US dollars per million tokens · most-agreed
Input
$0.04
Output
$0.15
Cache read
$0.004
Available from 7 hosts
| Host | In $/M | Out $/M | Cache rd | Cache wr | Context | Output | Retires |
|---|---|---|---|---|---|---|---|
| inception | $0.04 | $0.15 | $0.004 | — | 260K | 66K | — |
| nano-gpt · inception | $0.04 | $0.15 | $0.004 | — | 260K | 66K | — |
| openrouter · inception | $0.04 | $0.15 | $0.004 | — | 260K | 66K | — |
| vercel · inception | $0.04 | $0.15 | $0.004 | — | 260K | 66K | — |
| venice | $0.05 | $0.1875 | $0.005 | — | 260K | 66K | — |
| inception · inception | $0.20 | $0.75 | — | — | 260K | 66K | — |
| kilo · inception | $0.20 | $0.75 | $0.02 | — | 260K | 66K | — |
Price history
append-only observations · a listing writes a row only when its price moves
| Date | Host | Field | $/M |
|---|---|---|---|
| 2026-09-09 | venice | Cache read | $0.005 |
| 2026-09-10 | venice | Cache read | $0.025 |
| 2026-09-11 | venice | Cache read | $0.005 |
| 2026-09-09 | venice | Input | $0.05 |
| 2026-09-10 | venice | Input | $0.25 |
| 2026-09-11 | venice | Input | $0.05 |
| 2026-09-09 | venice | Output | $0.1875 |
| 2026-09-10 | venice | Output | $0.9375 |
| 2026-09-11 | venice | Output | $0.1875 |