modelbenchmark.io

grok-2

xAI · grok-2

Compare

Specification

most-agreed values

Context
131K
Max output
4K
Released
Knowledge cutoff
Retires
Open weights
Input
Output

Price

US dollars per million tokens · most-agreed

Input
$2.00
Output
$10.00

Quality

4 benchmarks · 1 source

Composite
32nd
percentile of 334 scored models
Rank
228 / 334
±0.05 sd
Evidence
4 × 1
one source, or fewer than 3 benchmarks
Effort range
one configuration only
reasoning39th1/4 bench
math31st3/7 bench

Every score

one row per source and configuration — nothing averaged away

BenchmarkScoreConfigurationSourceRun
MATH level 563.5 ±1.3defaultEpoch AI2025-02-17
GPQA diamond53.8 ±2.7defaultEpoch AI2025-01-27
FrontierMath-v10.3 ±0.5default · 2 runsEpoch AI2025-03-06
OTIS Mock AIME 2024-202511.5 ±3.4defaultEpoch AI2025-02-25

The leading figure for a benchmark is the median across its configurations, so one heroic high-effort run cannot set the number.

Available from 1 host

HostIn $/MOut $/MCache rdCache wrContextOutputRetires
vercel_ai_gateway · vercel-ai-gateway$2.00$10.00131K4K

More from xAI

most-hosted first