Compact GPT model for low-latency assistance and high-volume workloads
Specification
gemini listing
Context
1M
Max output
8K
Released
2024-12-06
Knowledge cutoff
—
Retires
—
Open weights
no
Input
text, image, audio
Output
text
Price
US dollars per million tokens · gemini list price
Input
free
Output
free
Cache read
$0.03
Long-context price tiers
the rate above each context threshold
| Above | Price | $/M |
|---|---|---|
| 128K | Input | free |
| 128K | Output | free |
Rate limits
as listed
Requests per minute
1,000
Tokens per minute
4,000,000
Available from 2 hosts
More from Google
most-hosted first