modelbenchmark.io

Grok 4 (xAI)

xAI · grok-4

Compare

Grok model for agentic tool use, reasoning, coding, and live assistance

Specification

xai listing

Context
256K
Max output
256K
Released
2025-07-09
Knowledge cutoff
2024-07
Retires
2026-05-15
Open weights
no
Input
text, image
Output
text

Price

US dollars per million tokens · xai list price

Input
$1.25
Output
$2.50
Cache read
$0.20
Cache write
$15.00

Long-context price tiers

the rate above each context threshold

AbovePrice$/M
200KCache read$0.40
200KInput$2.50
200KOutput$5.00

Quality

6 benchmarks · 2 sources

Composite
82nd
percentile of 334 scored models
Rank
61 / 334
±0.07 sd
Evidence
6 × 2
3 or more benchmarks, from 2 or more sources
Effort range
one configuration only
coding98th1/3 bench
reasoning91st2/4 bench
math62nd3/7 bench

Every score

one row per source and configuration — nothing averaged away

BenchmarkScoreConfigurationSourceRun
Aider polyglot79.6high · aider diffAider2025-07-11
GPQA diamond87.0 ±2.0defaultEpoch AI
OTIS Mock AIME 2024-202584.0 ±5.0defaultEpoch AI
Chess Puzzles28.0 ±4.5defaultEpoch AI2026-01-30
FrontierMath-v119.7 ±2.3defaultEpoch AI2025-11-13
FrontierMath-Tier-4-2025-07-011.0 ±0.0default · 2 runsEpoch AI2025-08-11

The leading figure for a benchmark is the median across its configurations, so one heroic high-effort run cannot set the number.

Available from 14 hosts

HostIn $/MOut $/MCache rdCache wrContextOutputRetires
xaivendor$1.25$2.50$0.20256K256K2026-05-15
jiekou$2.70$13.50256K8K
abacus$3.00$15.00256K16K
azure_ai$3.00$15.00131K131K
fastrouter · x-ai$3.00$15.00$0.75$15.00256K64K
helicone$3.00$15.00$0.75256K256K
llmgateway$3.00$15.00$0.75256K256K
llmgateway-providers$3.00$15.00$0.75256K256K
oci · oci$3.00$15.00128K128K
openrouter$3.00$15.00256K256K
poe$3.00$15.00$0.75256K128K
vercel_ai_gateway · vercel-ai-gateway$3.00$15.00256K256K
zenmux · x-ai$3.00$15.00$0.75256K64K
replicate · replicate$7.20$36.00

More from xAI

most-hosted first