modelbenchmark.io

Llama 4 Scout

Meta · llama-4-scout

Compare

Open multimodal Llama model for long-context analysis and efficient agents

Specification

most-agreed values

Context
328K
Max output
66K
Released
2025-09-05
Knowledge cutoff
2025-01
Retires
Open weights
yes
Input
text, image
Output
text

Price

US dollars per million tokens · most-agreed

Input
$0.085
Output
$0.46
Cache read
$0.0425

Quality

5 benchmarks · 2 sources

Composite
20th
percentile of 334 scored models
Rank
268 / 334
±0.05 sd
Evidence
5 × 2
3 or more benchmarks, from 2 or more sources
Effort range
one configuration only
coding2nd1/3 bench
reasoning36th1/4 bench
math29th3/7 bench

Every score

one row per source and configuration — nothing averaged away

BenchmarkScoreConfigurationSourceRun
MATH level 562.3 ±1.2defaultEpoch AI2025-04-08
GPQA diamond51.8 ±3.2defaultEpoch AI2025-04-08
FrontierMath-v10.0 ±0.0default · 2 runsEpoch AI2025-04-08
OTIS Mock AIME 2024-20257.8 ±3.0defaultEpoch AI2025-04-08
SWE-Bench verified9.1default · mini-SWE-agentSWE-bench

The leading figure for a benchmark is the median across its configurations, so one heroic high-effort run cannot set the number.

Available from 7 hosts

HostIn $/MOut $/MCache rdCache wrContextOutputRetires
vercelfreefree128K4K
helicone$0.08$0.30131K8K
nano-gpt · meta-llama$0.085$0.46$0.0425328K66K
kilo · meta-llama$0.10$0.30328K16K
openrouter$0.10$0.301.3M16K
openrouter · meta-llama$0.10$0.301.3M16K
vercel_ai_gateway · vercel-ai-gateway$0.10$0.30131K8K

More from Meta

most-hosted first