modelbenchmark.io

Claude Opus 4

anthropic-claude-opus-4

Compare

Flagship Claude model for deep reasoning, coding, and long-horizon agents

Specification

most-agreed values

Context
200K
Max output
32K
Released
2025-05-22
Knowledge cutoff
2025-03-31
Retires
Open weights
no
Input
text, image, pdf
Output
text

Price

US dollars per million tokens · most-agreed

Input
$15.00
Output
$75.00
Cache read
$1.50
Cache write
$18.75

Quality

7 benchmarks · 2 sources

Composite
70th
percentile of 334 scored models
Rank
101 / 334
±0.05 sd
Evidence
7 × 2
3 or more benchmarks, from 2 or more sources
Effort range
4.6
points between effort settings
coding78th2/3 bench
reasoning74th1/4 bench
math54th4/7 bench

By reasoning effort

every setting placed on the same scale as the leaderboard

SettingCompositeBenchmarks behind it
32k93rd1
16k73rd2
27k44th3

Every score

one row per source and configuration — nothing averaged away

BenchmarkScoreConfigurationSourceRun
Aider polyglot72.032k · aider diffAider2025-05-25
MATH level 585.0 ±1.0defaultEpoch AI2025-05-22
SWE-Bench verified70.7 ±2.1defaultEpoch AI2026-02-06
GPQA diamond76.3 ±3.016kEpoch AI2025-05-22
69.2 ±3.3defaultEpoch AI2025-05-22
Settings differ by 7.1 points on this benchmark.
OTIS Mock AIME 2024-202564.4 ±7.227kEpoch AI2025-05-28
60.0 ±7.416kEpoch AI2025-05-22
42.2 ±7.4defaultEpoch AI2025-05-22
Settings differ by 22.2 points on this benchmark.
FrontierMath-Tier-4-2025-07-012.1 ±0.027k · 2 runsEpoch AI2025-07-01
0.0 ±0.0default · 2 runsEpoch AI2025-07-01
Settings differ by 2.1 points on this benchmark.
FrontierMath-v14.1 ±1.227kEpoch AI2025-08-05
2.2 ±1.2default · 2 runsEpoch AI2025-07-04
Settings differ by 1.9 points on this benchmark.

The leading figure for a benchmark is the median across its configurations, so one heroic high-effort run cannot set the number.

Available from 1 host

HostIn $/MOut $/MCache rdCache wrContextOutputRetires
digitalocean$15.00$75.00$1.50$18.75200K32K