Efficient Mistral model for fast chat, extraction, and production assistants
Specification
most-agreed values
Context
262K
Max output
16K
Released
2026-03-17
Knowledge cutoff
2025-06
Retires
—
Open weights
yes
Input
text, image
Output
text
Price
US dollars per million tokens · most-agreed
Input
$0.40
Output
$1.40
Cache read
$0.20
Cache write
$0.15
Available from 6 hosts
| Host | In $/M | Out $/M | Cache rd | Cache wr | Context | Output | Retires |
|---|---|---|---|---|---|---|---|
| nvidia | free | free | — | — | 128K | 8K | — |
| pioneer | $0.15 | $0.60 | $0.015 | $0.15 | 32K | 32K | — |
| infomaniak | $0.25 | $0.93 | — | — | 256K | 256K | — |
| nano-gpt | $0.40 | $1.40 | $0.20 | — | 262K | 16K | — |
| nano-gpt · thinking | $0.40 | $1.40 | $0.20 | — | 262K | 16K | — |
| regolo-ai | $0.75 | $3.00 | — | — | 256K | 16K | — |
More from Mistral
most-hosted first