modelbenchmark.io

Gemini 3.1 Flash Lite

Google · gemini-3-1-flash-lite

Compare

Low-latency Gemini model for high-volume multimodal and agent workloads

Specification

google listing

Context
1M
Max output
66K
Released
2026-05-07
Knowledge cutoff
2025-01
Retires
2027-05-07
Open weights
no
Input
text, image, video, audio, pdf
Output
text

Price

US dollars per million tokens · google list price

Input
$0.25
Output
$1.50
Cache read
$0.025
Cache write
$0.0833

Batch and priority prices

US dollars per million tokens · per listing

PriceListing$/M
Batch inputvertex_ai-language-models$0.125
Batch outputvertex_ai-language-models$0.75
Priority inputvertex_ai-language-models$0.45
Priority outputvertex_ai-language-models$2.70

Batch input is 50% below the interactive input price at the same listing.

Rate limits

as listed

Requests per minute
15
Tokens per minute
250,000

Quality

4 benchmarks · 1 source

Composite
56th
percentile of 334 scored models
Rank
147 / 334
±0.08 sd
Evidence
4 × 1
one source, or fewer than 3 benchmarks
Effort range
7.2
points between effort settings
reasoning78th2/4 bench
math33rd2/7 bench

By reasoning effort

every setting placed on the same scale as the leaderboard

SettingCompositeBenchmarks behind it
high70th4
low57th4
minimal52nd4

Every score

one row per source and configuration — nothing averaged away

BenchmarkScoreConfigurationSourceRun
Chess Puzzles25.0 ±4.4lowEpoch AI2026-08-06
24.0 ±4.3minimalEpoch AI2026-08-06
20.0 ±4.0highEpoch AI2026-08-06
Settings differ by 5.0 points on this benchmark.
GPQA diamond81.8 ±2.7highEpoch AI2026-08-06
74.2 ±3.1lowEpoch AI2026-08-06
73.7 ±3.1minimalEpoch AI2026-08-06
Settings differ by 8.1 points on this benchmark.
OTIS Mock AIME 2024-202580.0 ±6.0highEpoch AI2026-08-06
44.4 ±7.5lowEpoch AI2026-08-06
37.8 ±7.3minimalEpoch AI2026-08-06
Settings differ by 42.2 points on this benchmark.
FrontierMath-Tiers-1-3-v227.7 ±2.7highEpoch AI2026-08-30
22.5 ±2.5lowEpoch AI2026-08-30
21.4 ±2.4minimalEpoch AI2026-08-28
Settings differ by 6.3 points on this benchmark.

The leading figure for a benchmark is the median across its configurations, so one heroic high-effort run cannot set the number.

Available from 32 hosts

HostIn $/MOut $/MCache rdCache wrContextOutputRetires
googlevendor$0.25$1.50$0.0251M66K
kenarifreefree1M66K
kilo$0.125$0.75$0.0125$0.04171M66K
302ai$0.25$1.501M66K
abacus$0.25$1.50$0.025$1.001M66K
aihubmix$0.25$1.50$0.025$1.001M66K
databricks
$0.25
$0.3125
$1.50
$1.875
$0.025
$0.0312
$0.31251M66K
deepinfra$0.25$1.501M
edenai$0.25$1.50$0.025$0.08331M66K
gemini$0.25$1.50$0.0251M66K
2026-05-25
2027-05-07
google-vertex$0.25$1.50$0.0251M66K
impossibl$0.25$1.50$0.0251M66K
llmgateway$0.25$1.50$0.025$0.08331M66K
llmgateway-providers · google-ai-studio$0.25$1.50$0.025$0.08331M66K
llmgateway-providers · google-vertex$0.25$1.50$0.025$0.08331M66K
merge-gateway$0.25$1.50$0.0251M66K
nano-gpt$0.25$1.50$0.025$0.08331M66K
nearai$0.25$1.50$0.0251M66K
neon$0.25$1.50$0.0251M66K
ofox$0.25$1.50$0.025$1.001M66K
openrouter$0.25$1.50$0.025$0.08331M66K
orcarouter$0.25$1.50$0.0251M66K
perplexity$0.25$1.50$0.025
pioneer$0.25$1.50$0.03$0.251M65K
poe$0.25$1.501M66K
requesty$0.25$1.50$0.025$0.08331M66K
sap-ai-core$0.25$1.50$0.0251M66K
vercel$0.25$1.50$0.031M65K
vertex_ai-language-models$0.25$1.50$0.0251M66K2027-05-07
vivgrid$0.25$1.50$0.025$1.001M66K
zenmux$0.25$1.50$0.025
1M
1.1M
66K
66K
cortecs$0.272$1.631$0.025$0.0821M66K

More from Google

most-hosted first