modelbenchmark.io

Deliberate o-series reasoner for hard math, coding, and multi-step analysis

Specification

openai listing

Context
200K
Max output
100K
Released
2025-04-16
Knowledge cutoff
2024-05
Retires
2026-12-11
Open weights
no
Input
text, image, pdf
Output
text

Price

US dollars per million tokens · openai list price

Input
$2.00
Output
$8.00
Cache read
$0.50
Cache write
free

Batch and priority prices

US dollars per million tokens · per listing

PriceListing$/M
Batch inputopenai$1.00
Batch outputopenai$4.00
Priority inputopenai$3.50
Priority outputopenai$14.00

Batch input is 50% below the interactive input price at the same listing.

Quality

11 benchmarks · 3 sources

Composite
76th
percentile of 334 scored models
Rank
80 / 334
±0.05 sd
Evidence
11 × 3
3 or more benchmarks, from 2 or more sources
Effort range
4.2
points between effort settings
coding74th2/3 bench
reasoning86th3/4 bench
math61st5/7 bench
knowledge70th1/1 bench

Every score

one row per source and configuration — nothing averaged away

BenchmarkScoreConfigurationSourceRun
MATH level 597.8 ±0.3defaultEpoch AI2025-04-16
Aider polyglot81.3high · aider diffAider2025-06-25
76.9default · aider diffAider2025-06-25
Settings differ by 4.4 points on this benchmark.
Chess Puzzles33.0 ±4.9default · 3 runsEpoch AI2026-08-07
GPQA diamond80.8 ±2.8default · 3 runsEpoch AI2026-08-07
OTIS Mock AIME 2024-202576.1 ±5.5default · 3 runsEpoch AI2026-08-07
SimpleQA Verified49.4 ±1.6defaultEpoch AI2026-08-27
Mystery Game Puzzles26.0 ±4.6default · 2 runsEpoch AI2026-08-27
SWE-Bench verified62.3 ±2.2defaultEpoch AI2026-02-12
58.4default · mini-SWE-agentSWE-bench
Sources and settings disagree by 3.9 points. Both figures stand.
FrontierMath-v110.9 ±1.7default · 6 runsEpoch AI2025-11-17
FrontierMath-Tier-4-2025-07-011.0 ±2.1default · 2 runsEpoch AI2025-07-01
FrontierMath-Tiers-1-3-v227.5 ±2.8default · 3 runsEpoch AI2026-08-27

The leading figure for a benchmark is the median across its configurations, so one heroic high-effort run cannot set the number.

Available from 21 hosts

HostIn $/MOut $/MCache rdCache wrContextOutputRetires
openaivendor$2.00$8.00$0.50200K100K2026-12-11
poe$1.80$7.20$0.45200K100K
302ai$2.00$8.00200K100K
abacus$2.00$8.00$0.50200K100K
azure
$2.00
$2.20
$8.00
$8.80
$0.50
$0.55
200K100K2026-11-19
azure-cognitive-services$2.00$8.00$0.50200K100K
cloudflare-ai-gateway$2.00$8.00$0.50200K100K
edenai$2.00$8.00$0.50200K100K
helicone$2.00$8.00$0.50200K100K
impossibl$2.00$8.00$0.50200K100K
kilo$2.00$8.00$0.50200K100K
llmgateway$2.00$8.00$0.50200K100K
llmgateway-providers$2.00$8.00$0.50200K100K
merge-gateway$2.00$8.00$0.50200K100K
nano-gpt$2.00$8.00$1.00200K100K
nearai$2.00$8.00$0.50200K100K
openrouter$2.00$8.00$0.50200K100K
vercel$2.00$8.00$0.50200K100K
vercel_ai_gateway · vercel-ai-gateway$2.00$8.00$0.50free200K100K
jiekou$10.00$40.00131K131K
anyapi200K100K

Price history

append-only observations · a listing writes a row only when its price moves

DateHostField$/M
2026-09-02openrouterCache read$0.50
2026-09-06openrouterCache read$0.50
2026-09-02openrouterInput$2.00
2026-09-06openrouterInput$2.00
2026-09-02openrouterOutput$8.00
2026-09-06openrouterOutput$8.00

Other listings of this model

same model, different key

More from OpenAI

most-hosted first