modelbenchmark.io

Gemini 3.7 Flash

Google · gemini-3-7-flash

Compare

High-efficiency Gemini model for agentic workflows, coding, and multimodal reasoning

Specification

google listing

Context
1M
Max output
66K
Released
2026-08-13
Knowledge cutoff
2026-03
Retires
Open weights
no
Input
text, image, video, audio, pdf
Output
text

Price

US dollars per million tokens · google list price

Input
$0.75
Output
$3.75
Cache read
$0.075
Cache write
$0.0417

Batch and priority prices

US dollars per million tokens · per listing

PriceListing$/M
Batch inputvertex_ai$0.375
Batch outputvertex_ai$1.875
Priority inputvertex_ai$1.35
Priority outputvertex_ai$6.75

Batch input is 50% below the interactive input price at the same listing.

Rate limits

as listed

Requests per minute
2,000
Tokens per minute
800,000

Quality

8 benchmarks · 2 sources

Composite
94th
percentile of 334 scored models
Rank
22 / 334
±0.08 sd
Evidence
8 × 2
3 or more benchmarks, from 2 or more sources
Effort range
one configuration only
reasoning97th3/4 bench
math89th3/7 bench
knowledge92nd1/1 bench
general86th1/1 bench

Every score

one row per source and configuration — nothing averaged away

BenchmarkScoreConfigurationSourceRun
Chess Puzzles47.0 ±5.0highEpoch AI2026-08-14
SimpleQA Verified69.2 ±1.5highEpoch AI2026-08-27
GPQA diamond94.8 ±1.3highEpoch AI2026-08-14
OTIS Mock AIME 2024-202597.2 ±1.9highEpoch AI2026-08-14
Mystery Game Puzzles37.0 ±4.9highEpoch AI2026-08-14
FrontierMath-Tiers-1-3-v271.6 ±2.7highEpoch AI2026-08-14
LiveBench79.9high · 23 runsLiveBench2026-06-25
FrontierMath-Tier-4-v236.6 ±7.6highEpoch AI2026-08-14

The leading figure for a benchmark is the median across its configurations, so one heroic high-effort run cannot set the number.

Available from 30 hosts

HostIn $/MOut $/MCache rdCache wrContextOutputRetires
googlevendor$0.75$3.75$0.0751M66K
kenarifreefree1M66K
302ai$0.75$3.751M66K
abacus$0.75$3.75$0.0751M66K
aihubmix$0.75$3.75$0.0751M66K
cortecs$0.75$3.75$0.075$0.0381M66K
crossmodel$0.75$3.75$0.075$0.751M66K
deepinfra$0.75$3.751M
edenai$0.75$3.75$0.075$0.04171M66K
gemini$0.75$3.75$0.0751M66K
github-copilot$0.75$3.75$0.0751M64K
google-vertex$0.75$3.75$0.0751M66K
kilo$0.75$3.75$0.075$0.04171M66K
llmgateway$0.75$3.75$0.075$0.08331M66K
llmgateway-providers · google-ai-studio$0.75$3.75$0.075$0.08331M66K
llmgateway-providers · google-vertex$0.75$3.75$0.075$0.08331M66K
merge-gateway$0.75$3.75$0.0751M66K
nano-gpt$0.75$3.75$0.075$0.04171M66K
ofox$0.75$3.75$0.075$0.04151M66K
openrouter$0.75$3.75$0.075$0.04171M66K
opper · vertexai$0.75$3.75$0.0751M66K
perplexity$0.75$3.75$0.075
requesty$0.75$3.75$0.0751M66K
vercel$0.75$3.75$0.0751M66K
vertex_ai$0.75$3.75$0.0751M66K
vertex_ai-language-models$0.75$3.75$0.0751M66K
vivgrid$0.75$3.75$0.0751M66K
venice$0.9375$4.6875$0.09381M66K
opencode$1.50$7.50$0.151M66K
databricks1M66K

Price history

append-only observations · a listing writes a row only when its price moves

DateHostField$/M
2026-09-02nano-gptCache read$0.0375
2026-09-02nano-gptCache read$0.075
2026-09-02openrouterCache read$0.075
2026-09-06openrouterCache read$0.075
2026-09-02nano-gptCache write$0.0208
2026-09-02nano-gptCache write$0.075
2026-09-09nano-gptCache write$0.0417
2026-09-02nano-gptInput$0.375
2026-09-02nano-gptInput$0.75
2026-09-02openrouterInput$0.75
2026-09-06openrouterInput$0.75
2026-09-02nano-gptOutput$1.875
2026-09-02nano-gptOutput$3.75
2026-09-02openrouterOutput$3.75
2026-09-06openrouterOutput$3.75

More from Google

most-hosted first