modelbenchmark.io

Anthropic: Claude Opus 4

Anthropic · claude-4-opus

Compare

Claude Opus 4 is benchmarked as the world’s best coding model, at time of release, bringing sustained performance on complex, long-running tasks and agent workflows. It sets new benchmarks in...

Specification

anthropic listing

Context
200K
Max output
32K
Released
Knowledge cutoff
2025-01-31
Retires
2026-06-15
Open weights
Input
image, text, file
Output
text

Price

US dollars per million tokens · anthropic list price

Input
$15.00
Output
$75.00
Cache read
$1.50
Cache write
$18.75

Quality

1 benchmark · 1 source

Composite
85th
percentile of 334 scored models
Rank
52 / 334
±0.10 sd
Evidence
1 × 1
too little evidence to place with confidence
Effort range
one configuration only
coding67th1/3 bench

Every score

one row per source and configuration — nothing averaged away

BenchmarkScoreConfigurationSourceRun
SWE-Bench verified70.4default · Tools · 2 runsSWE-bench

The leading figure for a benchmark is the median across its configurations, so one heroic high-effort run cannot set the number.

Available from 4 hosts

HostIn $/MOut $/MCache rdCache wrContextOutputRetires
anthropicvendor$15.00$75.00$1.50$18.75200K32K2026-06-15
snowflake$5.00$25.00$0.50200K16K
vercel_ai_gateway · vercel-ai-gateway$15.00$75.00$1.50$18.75200K32K
deepinfra$16.50$82.50200K200K

More from Anthropic

most-hosted first