modelbenchmark.io

Anthropic: Claude Sonnet 4

Anthropic · claude-4-sonnet

Compare

Claude Sonnet 4 significantly enhances the capabilities of its predecessor, Sonnet 3.7, excelling in both coding and reasoning tasks with improved precision and controllability. Achieving state-of-the-art performance on SWE-bench (72.7%),...

Specification

anthropic listing

Context
1M
Max output
64K
Released
Knowledge cutoff
2025-01-31
Retires
2026-06-15
Open weights
Input
image, text, file
Output
text

Price

US dollars per million tokens · anthropic list price

Input
$3.00
Output
$15.00
Cache read
$0.30
Cache write
$3.75

Long-context price tiers

the rate above each context threshold

AbovePrice$/M
200KCache read$0.60
200KCache write$7.50
200KInput$6.00
200KOutput$22.50

Quality

1 benchmark · 1 source

Composite
86th
percentile of 334 scored models
Rank
49 / 334
±0.10 sd
Evidence
1 × 1
too little evidence to place with confidence
Effort range
one configuration only
coding68th1/3 bench

Every score

one row per source and configuration — nothing averaged away

BenchmarkScoreConfigurationSourceRun
SWE-Bench verified71.0default · SWE-agent · 8 runsSWE-bench

The leading figure for a benchmark is the median across its configurations, so one heroic high-effort run cannot set the number.

Available from 6 hosts

HostIn $/MOut $/MCache rdCache wrContextOutputRetires
anthropicvendor$3.00$15.00$0.30$3.751M64K2026-06-15
replicate · replicate$3.00$15.00
snowflake$3.00$15.00$0.30200K16K
vercel_ai_gateway · vercel-ai-gateway$3.00$15.00$0.30$3.75200K64K
deepinfra$3.30$16.50200K200K
heroku · heroku200K8K

More from Anthropic

most-hosted first