modelbenchmark.io

Cerebras-Llama-4-Scout-17B-16E-Instruct

Meta · llama-4-scout-17b-16e-instruct

Compare

Open multimodal Llama model for long-context analysis and efficient agents

Specification

most-agreed values

Context
128K
Max output
4K
Released
2025-04-05
Knowledge cutoff
2025-01
Retires
2026-07-17
Open weights
yes
Input
text
Output
text

Price

US dollars per million tokens · most-agreed

Input
free
Output
free

Rate limits

Requests per minute
300

Available from 17 hosts

HostIn $/MOut $/MCache rdCache wrContextOutputRetires
llamafreefree128K4K
lambda_ai$0.05$0.1016K8K
nscale$0.09$0.29
deepinfra$0.10$0.30328K328K
deepinfra · meta-llama$0.10$0.30328K16K
groq$0.11$0.34131K8K2026-07-17
wandb · wandb$0.17$0.6664K64K
novita$0.18$0.59131K131K
novita-ai · meta-llama$0.18$0.59131K131K
together_ai$0.18$0.592026-02-06
azure$0.20$0.78128K8K
azure_ai$0.20$0.7810M16K
azure-cognitive-services$0.20$0.78128K8K
cloudflare$0.27$0.85131K131K
cloudflare-workers-ai · cf$0.27$0.85131K16K
sambanova · sambanova$0.40$0.708K8K2025-06-19
oci · oci$0.72$0.7210.5M8K

More from Meta

most-hosted first