modelbenchmark.io

Gemma 3 4B IT

Google · gemma-3-4b-it

Compare

Open Gemma instruction model for efficient chat and self-hosted deployments

Specification

most-agreed values

Context
131K
Max output
131K
Released
2024-01-01
Knowledge cutoff
2024-08
Retires
Open weights
yes
Input
text, pdf
Output
text

Price

US dollars per million tokens · most-agreed

Input
$0.05
Output
$0.10
Cache read
$0.1003

Quality

3 benchmarks · 1 source

Composite
10th
percentile of 334 scored models
Rank
300 / 334
±0.10 sd
Evidence
3 × 1
one source, or fewer than 3 benchmarks
Effort range
one configuration only
reasoning6th2/4 bench
math10th1/7 bench

Every score

one row per source and configuration — nothing averaged away

BenchmarkScoreConfigurationSourceRun
Chess Puzzles0.0 ±0.0defaultEpoch AI2026-08-28
OTIS Mock AIME 2024-20257.5 ±2.8defaultEpoch AI2026-08-28
GPQA diamond23.2 ±1.9defaultEpoch AI2026-08-28

The leading figure for a benchmark is the median across its configurations, so one heroic high-effort run cannot set the number.

Available from 10 hosts

HostIn $/MOut $/MCache rdCache wrContextOutputRetires
nvidiafreefree131K16K
amazon-bedrock$0.04$0.08131K4K
bedrock_converse$0.04$0.08128K8K
merge-gateway$0.04$0.08128K8K
deepinfra$0.05$0.10131K131K
edenai
$0.05
$0.04
$0.10
$0.08
131K
128K
131K
huggingface$0.05$0.10131K131K
kilo$0.05$0.10131K16K
openrouter$0.05$0.10131K16K
nano-gpt · unsloth$0.2006$0.2006$0.1003128K8K

Price history

append-only observations · a listing writes a row only when its price moves

DateHostField$/M
2026-09-02deepinfraInput$0.05
2026-09-12deepinfraInput$0.05
2026-09-02openrouterInput$0.05
2026-09-06openrouterInput$0.05
2026-09-02deepinfraOutput$0.10
2026-09-12deepinfraOutput$0.10
2026-09-02openrouterOutput$0.10
2026-09-06openrouterOutput$0.10

More from Google

most-hosted first