modelbenchmark.io

Claude Sonnet 4

anthropic-claude-sonnet-4

Compare

Balanced Claude model for coding, analysis, agent workflows, and cost control

Specification

most-agreed values

Context
1M
Max output
64K
Released
2025-05-22
Knowledge cutoff
2025-03-31
Retires
Open weights
no
Input
text, image, pdf
Output
text

Price

US dollars per million tokens · most-agreed

Input
$3.00
Output
$15.00
Cache read
$0.30
Cache write
$3.75

Quality

7 benchmarks · 3 sources

Composite
66th
percentile of 334 scored models
Rank
113 / 334
±0.05 sd
Evidence
7 × 3
3 or more benchmarks, from 2 or more sources
Effort range
11.6
points between effort settings
coding64th2/3 bench
reasoning80th1/4 bench
math53rd4/7 bench

By reasoning effort

every setting placed on the same scale as the leaderboard

SettingCompositeBenchmarks behind it
32k82nd3
16k69th2
59k64th3

Every score

one row per source and configuration — nothing averaged away

BenchmarkScoreConfigurationSourceRun
MATH level 584.4 ±1.0defaultEpoch AI2025-05-22
Aider polyglot61.332k · aider diffAider2025-05-24
GPQA diamond78.3 ±2.932kEpoch AI2025-05-22
77.8 ±2.759kEpoch AI2025-05-26
75.8 ±3.116kEpoch AI2025-05-22
66.7 ±3.4defaultEpoch AI2025-05-22
Settings differ by 11.6 points on this benchmark.
SWE-Bench verified65.9default · Harness AI · 2 runsSWE-bench
OTIS Mock AIME 2024-202571.1 ±6.832kEpoch AI2025-05-22
68.9 ±7.059kEpoch AI2025-05-23
53.3 ±7.516kEpoch AI2025-05-22
28.9 ±6.8defaultEpoch AI2025-05-22
Settings differ by 42.2 points on this benchmark.
FrontierMath-Tier-4-2025-07-010.0 ±0.0default · 2 runsEpoch AI2025-07-01
0.0 ±0.059kEpoch AI2025-07-01
FrontierMath-v12.1 ±1.2default · 2 runsEpoch AI2025-07-04

The leading figure for a benchmark is the median across its configurations, so one heroic high-effort run cannot set the number.

Available from 1 host

HostIn $/MOut $/MCache rdCache wrContextOutputRetires
digitalocean$3.00$15.00$0.30$3.751M64K