modelbenchmark.io

GLM 4.7 Flash Original

Zhipu · glm-4-7-flash-original

Compare

GLM-4.7-Flash is a lightweight 30B model optimized for coding and agentic tasks. Balances high performance with efficiency, perfect for local deployment.

Specification

most-agreed values

Context
200K
Max output
128K
Released
2026-01-19
Knowledge cutoff
Retires
Open weights
yes
Input
text
Output
text

Price

US dollars per million tokens · most-agreed

Input
$0.07
Output
$0.40
Cache read
$0.035

Available from 2 hosts

HostIn $/MOut $/MCache rdCache wrContextOutputRetires
nano-gpt · thinking$0.07$0.40$0.035200K128K
nano-gpt · z-ai$0.07$0.40$0.035200K128K

More from Zhipu

most-hosted first