modelbenchmark.io

GLM 4.1V Thinking Flash

Zhipu · glm-4-1v-thinking-flash

Compare

Compact GPT model for low-latency assistance and high-volume workloads

Specification

most-agreed values

Context
64K
Max output
8K
Released
2025-07-09
Knowledge cutoff
Retires
Open weights
no
Input
text, image
Output
text

Price

US dollars per million tokens · most-agreed

Input
$0.30
Output
$0.30
Cache read
$0.15

Available from 1 host

HostIn $/MOut $/MCache rdCache wrContextOutputRetires
nano-gpt$0.30$0.30$0.1564K8K

More from Zhipu

most-hosted first