GLM-5.3-Prime is the high-speed variant of Z.ai's GLM-5.3, inheriting its full capabilities while delivering 1.5–2× the output throughput through inference acceleration. It supports text input and output with a 1M-token...
Specification
most-agreed values
Context
1M
Max output
131K
Released
2026-09-23
Knowledge cutoff
—
Retires
—
Open weights
no
Input
text
Output
text
Price
US dollars per million tokens · most-agreed
Input
$2.80
Output
$8.80
Cache read
$0.56
Available from 3 hosts
| Host | In $/M | Out $/M | Cache rd | Cache wr | Context | Output | Retires |
|---|---|---|---|---|---|---|---|
| kilo · z-ai | $2.80 | $8.80 | $0.56 | — | 1M | 131K | — |
| openrouter | $2.80 | $8.80 | $0.56 | — | 1M | 131K | — |
| openrouter · z-ai | $2.80 | $8.80 | $0.56 | — | 1M | 131K | — |
More from Zhipu
most-hosted first