Low-latency Gemini model for high-volume multimodal and agent workloads
Specification
google listing
Context
1M
Max output
66K
Released
2026-05-07
Knowledge cutoff
2025-01
Retires
2027-05-07
Open weights
no
Input
text, image, video, audio, pdf
Output
text
Price
US dollars per million tokens · google list price
Input
$0.25
Output
$1.50
Cache read
$0.025
Cache write
$0.0833
Batch and priority prices
US dollars per million tokens · per listing
| Price | Listing | $/M |
|---|---|---|
| Batch input | vertex_ai-language-models | $0.125 |
| Batch output | vertex_ai-language-models | $0.75 |
| Priority input | vertex_ai-language-models | $0.45 |
| Priority output | vertex_ai-language-models | $2.70 |
Batch input is 50% below the interactive input price at the same listing.
Rate limits
as listed
Requests per minute
15
Tokens per minute
250,000
Quality
4 benchmarks · 1 source
Composite
56th
percentile of 334 scored models
Rank
147 / 334
±0.08 sd
Evidence
4 × 1
one source, or fewer than 3 benchmarks
Effort range
7.2
points between effort settings
reasoning78th2/4 bench
math33rd2/7 bench
By reasoning effort
every setting placed on the same scale as the leaderboard
| Setting | Composite | Benchmarks behind it |
|---|---|---|
| high | 70th | 4 |
| low | 57th | 4 |
| minimal | 52nd | 4 |
Every score
one row per source and configuration — nothing averaged away
| Benchmark | Score | Configuration | Source | Run |
|---|---|---|---|---|
| Chess Puzzles | 25.0 ±4.4 | low | Epoch AI | 2026-08-06 |
| ↳ | 24.0 ±4.3 | minimal | Epoch AI | 2026-08-06 |
| ↳ | 20.0 ±4.0 | high | Epoch AI | 2026-08-06 |
| Settings differ by 5.0 points on this benchmark. | ||||
| GPQA diamond | 81.8 ±2.7 | high | Epoch AI | 2026-08-06 |
| ↳ | 74.2 ±3.1 | low | Epoch AI | 2026-08-06 |
| ↳ | 73.7 ±3.1 | minimal | Epoch AI | 2026-08-06 |
| Settings differ by 8.1 points on this benchmark. | ||||
| OTIS Mock AIME 2024-2025 | 80.0 ±6.0 | high | Epoch AI | 2026-08-06 |
| ↳ | 44.4 ±7.5 | low | Epoch AI | 2026-08-06 |
| ↳ | 37.8 ±7.3 | minimal | Epoch AI | 2026-08-06 |
| Settings differ by 42.2 points on this benchmark. | ||||
| FrontierMath-Tiers-1-3-v2 | 27.7 ±2.7 | high | Epoch AI | 2026-08-30 |
| ↳ | 22.5 ±2.5 | low | Epoch AI | 2026-08-30 |
| ↳ | 21.4 ±2.4 | minimal | Epoch AI | 2026-08-28 |
| Settings differ by 6.3 points on this benchmark. | ||||
The leading figure for a benchmark is the median across its configurations, so one heroic high-effort run cannot set the number.
Available from 32 hosts
| Host | In $/M | Out $/M | Cache rd | Cache wr | Context | Output | Retires |
|---|---|---|---|---|---|---|---|
| googlevendor | $0.25 | $1.50 | $0.025 | — | 1M | 66K | — |
| kenari | free | free | — | — | 1M | 66K | — |
| kilo | $0.125 | $0.75 | $0.0125 | $0.0417 | 1M | 66K | — |
| 302ai | $0.25 | $1.50 | — | — | 1M | 66K | — |
| abacus | $0.25 | $1.50 | $0.025 | $1.00 | 1M | 66K | — |
| aihubmix | $0.25 | $1.50 | $0.025 | $1.00 | 1M | 66K | — |
| databricks | $0.25 $0.3125 | $1.50 $1.875 | $0.025 $0.0312 | $0.3125 | 1M | 66K | — |
| deepinfra | $0.25 | $1.50 | — | — | 1M | — | — |
| edenai | $0.25 | $1.50 | $0.025 | $0.0833 | 1M | 66K | — |
| gemini | $0.25 | $1.50 | $0.025 | — | 1M | 66K | 2026-05-25 2027-05-07 |
| google-vertex | $0.25 | $1.50 | $0.025 | — | 1M | 66K | — |
| impossibl | $0.25 | $1.50 | $0.025 | — | 1M | 66K | — |
| llmgateway | $0.25 | $1.50 | $0.025 | $0.0833 | 1M | 66K | — |
| llmgateway-providers · google-ai-studio | $0.25 | $1.50 | $0.025 | $0.0833 | 1M | 66K | — |
| llmgateway-providers · google-vertex | $0.25 | $1.50 | $0.025 | $0.0833 | 1M | 66K | — |
| merge-gateway | $0.25 | $1.50 | $0.025 | — | 1M | 66K | — |
| nano-gpt | $0.25 | $1.50 | $0.025 | $0.0833 | 1M | 66K | — |
| nearai | $0.25 | $1.50 | $0.025 | — | 1M | 66K | — |
| neon | $0.25 | $1.50 | $0.025 | — | 1M | 66K | — |
| ofox | $0.25 | $1.50 | $0.025 | $1.00 | 1M | 66K | — |
| openrouter | $0.25 | $1.50 | $0.025 | $0.0833 | 1M | 66K | — |
| orcarouter | $0.25 | $1.50 | $0.025 | — | 1M | 66K | — |
| perplexity | $0.25 | $1.50 | $0.025 | — | — | — | — |
| pioneer | $0.25 | $1.50 | $0.03 | $0.25 | 1M | 65K | — |
| poe | $0.25 | $1.50 | — | — | 1M | 66K | — |
| requesty | $0.25 | $1.50 | $0.025 | $0.0833 | 1M | 66K | — |
| sap-ai-core | $0.25 | $1.50 | $0.025 | — | 1M | 66K | — |
| vercel | $0.25 | $1.50 | $0.03 | — | 1M | 65K | — |
| vertex_ai-language-models | $0.25 | $1.50 | $0.025 | — | 1M | 66K | 2027-05-07 |
| vivgrid | $0.25 | $1.50 | $0.025 | $1.00 | 1M | 66K | — |
| zenmux | $0.25 | $1.50 | $0.025 | — | 1M 1.1M | 66K 66K | — |
| cortecs | $0.272 | $1.631 | $0.025 | $0.082 | 1M | 66K | — |
More from Google
most-hosted first