modelbenchmark.io

Grok 4.20

xAI · grok-4-20

Compare

Grok model for agentic tool use, reasoning, coding, and live assistance

Specification

xai listing

Context
1M
Max output
1M
Released
2026-03-31
Knowledge cutoff
2025-09-01
Retires
Open weights
no
Input
text, image
Output
text

Price

US dollars per million tokens · xai list price

Input
$1.25
Output
$2.50
Cache read
$0.20

Long-context price tiers

the rate above each context threshold

AbovePrice$/M
200KCache read$0.40
200KInput$2.50
200KOutput$5.00

Quality

6 benchmarks · 1 source

Composite
69th
percentile of 334 scored models
Rank
104 / 334
±0.07 sd
Evidence
6 × 1
one source, or fewer than 3 benchmarks
Effort range
one configuration only
reasoning89th2/4 bench
math66th3/7 bench
knowledge25th1/1 bench

Every score

one row per source and configuration — nothing averaged away

BenchmarkScoreConfigurationSourceRun
OTIS Mock AIME 2024-202592.2 ±3.3defaultEpoch AI2026-07-13
GPQA diamond89.3 ±1.9defaultEpoch AI2026-07-13
Chess Puzzles24.0 ±4.3defaultEpoch AI2026-07-13
FrontierMath-Tiers-1-3-v244.9 ±3.0defaultEpoch AI2026-07-13
SimpleQA Verified30.2 ±1.5defaultEpoch AI2026-08-27
FrontierMath-Tier-4-v217.1 ±5.9defaultEpoch AI2026-07-13

The leading figure for a benchmark is the median across its configurations, so one heroic high-effort run cannot set the number.

Available from 8 hosts

HostIn $/MOut $/MCache rdCache wrContextOutputRetires
xaivendor$1.25$2.50$0.201M1M
kilo · x-ai$1.25$2.50$0.202M1.8M
openrouter$1.25$2.50$0.201M1M
openrouter · x-ai$1.25$2.50$0.202M1.8M
venice$1.42$2.83$0.232M128K
nano-gpt · x-ai$2.00$6.00$1.002M131K
oci · oci$3.00$15.00131K131K
ofox · x-ai$4.00$12.00$0.402M128K

More from xAI

most-hosted first