modelbenchmark.io

Claude 3.7 Sonnet

anthropic-claude-3-7-sonnet

Compare

Legacy model retained for compatibility with older integrations

Specification

most-agreed values

Context
200K
Max output
64K
Released
2025-02-24
Knowledge cutoff
2024-11
Retires
Open weights
no
Input
text, image
Output
text

Price

US dollars per million tokens · most-agreed

Input
$3.00
Output
$15.00
Cache read
$0.30
Cache write
$3.75

Quality

7 benchmarks · 3 sources

Composite
64th
percentile of 334 scored models
Rank
122 / 334
±0.05 sd
Evidence
7 × 3
3 or more benchmarks, from 2 or more sources
Effort range
12.5
points between effort settings
coding60th2/3 bench
reasoning80th1/4 bench
math50th4/7 bench

By reasoning effort

every setting placed on the same scale as the leaderboard

SettingCompositeBenchmarks behind it
32k71st5
16k63rd4
64k62nd5

Every score

one row per source and configuration — nothing averaged away

BenchmarkScoreConfigurationSourceRun
MATH level 591.2 ±0.864kEpoch AI2025-03-13
90.0 ±0.832kEpoch AI2025-03-12
86.2 ±0.916kEpoch AI2025-02-26
68.2 ±1.1defaultEpoch AI2025-02-24
Settings differ by 23.0 points on this benchmark.
Aider polyglot64.932k · aider diffAider2025-02-24
GPQA diamond78.5 ±2.764kEpoch AI2025-05-26
76.8 ±3.016kEpoch AI2025-02-26
76.8 ±3.032kEpoch AI2025-03-10
66.0 ±2.7defaultEpoch AI2025-02-24
Settings differ by 12.5 points on this benchmark.
SWE-Bench verified61.0 ±2.2defaultEpoch AI2026-02-04
59.6default · Aime-coder v1 · 2 runsSWE-bench
Sources and settings disagree by 1.4 points. Both figures stand.
OTIS Mock AIME 2024-202557.8 ±7.464kEpoch AI2025-03-13
53.3 ±7.532kEpoch AI2025-03-12
46.7 ±7.516kEpoch AI2025-02-26
21.9 ±4.8defaultEpoch AI2025-02-25
Settings differ by 35.9 points on this benchmark.
FrontierMath-Tier-4-2025-07-010.0 ±0.064kEpoch AI2025-07-01
FrontierMath-v12.1 ±1.216k · 2 runsEpoch AI2025-03-06
1.7 ±1.132k · 2 runsEpoch AI2025-03-13
1.6 ±0.0default · 2 runsEpoch AI2025-03-06
1.6 ±1.064k · 2 runsEpoch AI2025-03-13
Settings differ by 0.5 points on this benchmark.

The leading figure for a benchmark is the median across its configurations, so one heroic high-effort run cannot set the number.

Available from 3 hosts

HostIn $/MOut $/MCache rdCache wrContextOutputRetires
digitalocean$3.00$15.00$0.30$3.75200K64K
gradient_ai · gradient-ai$3.00$15.00200K1K
sap-ai-core$3.00$15.00$0.30$3.75200K64K

Other listings of this model

this page is one of them

The corpus files Claude 3.7 Sonnet under several keys. This page is the anthropic-claude-3-7-sonnet listing. The full record — every host, every price and every benchmark score — is on the main Claude 3.7 Sonnet page.