modelbenchmark.io

claude-3-sonnet

Anthropic · claude-3-sonnet

Compare

Specification

most-agreed values

Context
200K
Max output
4K
Released
Knowledge cutoff
Retires
2026-07-30
Open weights
Input
Output

Price

US dollars per million tokens · most-agreed

Input
$3.00
Output
$15.00
Cache read
$0.30
Cache write
$3.75

Quality

3 benchmarks · 1 source

Composite
13th
percentile of 334 scored models
Rank
291 / 334
±0.06 sd
Evidence
3 × 1
one source, or fewer than 3 benchmarks
Effort range
one configuration only
reasoning17th1/4 bench
math9th2/7 bench

Every score

one row per source and configuration — nothing averaged away

BenchmarkScoreConfigurationSourceRun
GPQA diamond40.6 ±2.4defaultEpoch AI2025-01-27
MATH level 518.2 ±0.8defaultEpoch AI2025-01-27
OTIS Mock AIME 2024-20252.5 ±1.9defaultEpoch AI2025-02-25

The leading figure for a benchmark is the median across its configurations, so one heroic high-effort run cannot set the number.

Available from 5 hosts

HostIn $/MOut $/MCache rdCache wrContextOutputRetires
bedrock$3.00$15.00$0.30$3.75200K4K2026-07-30
bedrock · apac$3.00$15.00$0.30$3.75200K4K2026-07-30
bedrock · eu$3.00$15.00$0.30$3.75200K4K2026-07-30
bedrock · us$3.00$15.00$0.30$3.75200K4K2026-07-30
vertex_ai-anthropic_models$3.00$15.00200K4K

More from Anthropic

most-hosted first