Compact GPT model for low-latency assistance and high-volume workloads
2 figures withheld: the value failed a range check.
Specification
most-agreed values
Context
128K
Max output
16K
Released
2025-07-26
Knowledge cutoff
2024-12
Retires
—
Open weights
yes
Input
text
Output
text
Price
US dollars per million tokens · most-agreed
Input
$0.17
Output
$0.68
Cache read
$0.085
Available from 4 hosts
| Host | In $/M | Out $/M | Cache rd | Cache wr | Context | Output | Retires |
|---|---|---|---|---|---|---|---|
| nvidia | free | free | — | — | 131K | 8K | — |
| azure_ai | $0.075 | $0.30 | — | — | 131K | 4K | — |
| nano-gpt | $0.17 | $0.68 | $0.085 | — | 128K | 16K | — |
| wandb · wandb | — | — | — | — | 128K | 128K | — |
Other listings of this model
same model, different key
More from Microsoft
most-hosted first