Low-latency Gemini model for high-volume multimodal and agent workloads
Specification
gemini listing
Context
1M
Max output
8K
Released
2025-08-05
Knowledge cutoff
2024-11
Retires
2026-06-01
Open weights
no
Input
text, image, audio, video
Output
text
Price
US dollars per million tokens · gemini list price
Input
$0.075
Output
$0.30
Cache read
$0.0187
Batch and priority prices
US dollars per million tokens · per listing
| Price | Listing | $/M |
|---|---|---|
| Batch input | vertex_ai-language-models | $0.0375 |
| Batch output | vertex_ai-language-models | $0.15 |
Batch input is 50% below the interactive input price at the same listing.
Rate limits
as listed
Requests per minute
4,000
Tokens per minute
4,000,000
Available from 6 hosts
| Host | In $/M | Out $/M | Cache rd | Cache wr | Context | Output | Retires |
|---|---|---|---|---|---|---|---|
| geminivendor | $0.075 | $0.30 | $0.0187 | — | 1M | 8K | 2026-06-01 |
| poe | $0.052 | $0.21 | — | — | 990K | 8K | — |
| 302ai | $0.075 | $0.30 | — | — | 2M | 8K | — |
| vercel_ai_gateway · vercel-ai-gateway | $0.075 | $0.30 | — | — | 1M | 8K | 2026-06-01 |
| vertex_ai-language-models | $0.075 | $0.30 | $0.0187 | — | 1M | 8K | 2026-06-01 |
| qiniu-ai | — | — | — | — | 1M | 8K | — |
More from Google
most-hosted first