modelbenchmark.io

Gemma 4 31B

Google · gemma-4-31b

Compare

Largest Gemma 4 instruction model for open, self-hosted chat and reasoning

Specification

most-agreed values

Context
262K
Max output
16K
Released
2026-04-02
Knowledge cutoff
Retires
Open weights
yes
Input
text, image
Output
text

Price

US dollars per million tokens · most-agreed

Input
$0.144
Output
$0.42
Cache read
$0.0144

Quality

4 benchmarks · 1 source

Composite
43rd
percentile of 334 scored models
Rank
191 / 334
±0.08 sd
Evidence
4 × 1
one source, or fewer than 3 benchmarks
Effort range
one configuration only
reasoning56th2/4 bench
math82nd1/7 bench
knowledge3rd1/1 bench

Every score

one row per source and configuration — nothing averaged away

BenchmarkScoreConfigurationSourceRun
OTIS Mock AIME 2024-202573.3 ±6.7minimalEpoch AI2026-08-06
GPQA diamond75.8 ±3.1minimalEpoch AI2026-08-06
Chess Puzzles5.0 ±2.2minimalEpoch AI2026-08-06
SimpleQA Verified10.4 ±1.0defaultEpoch AI2026-08-27

The leading figure for a benchmark is the median across its configurations, so one heroic high-effort run cannot set the number.

Available from 5 hosts

HostIn $/MOut $/MCache rdCache wrContextOutputRetires
poefreefree262K8K
amazon-bedrock$0.14$0.40262K33K
bedrock_mantle · bedrock-mantle
$0.14
$0.168
$0.40
$0.48
256K256K
neuralwatt$0.144$0.42$0.0144262K16K
cerebras$0.99$1.49131K41K

Price history

append-only observations · a listing writes a row only when its price moves

DateHostField$/M
2026-09-02bedrock_mantle · bedrock-mantleInput$0.14
2026-09-09bedrock_mantle · bedrock-mantleInput$0.14
2026-09-02cerebrasInput$0.99
2026-09-07cerebrasInput$0.99
2026-09-02bedrock_mantle · bedrock-mantleOutput$0.40
2026-09-09bedrock_mantle · bedrock-mantleOutput$0.40
2026-09-02cerebrasOutput$1.49
2026-09-07cerebrasOutput$1.49

Other listings of this model

same model, different key

More from Google

most-hosted first