modelbenchmark.io

Llama 4 Maverick 17B 128E Instruct FP8

Meta · llama-4-maverick-17b-128e-instruct-fp8

Compare

Open multimodal Llama for strong reasoning with efficient everyday serving

Specification

most-agreed values

Context
131K
Max output
8K
Released
2025-04-05
Knowledge cutoff
2024-08
Retires
2026-03-31
Open weights
yes
Input
text, image
Output
text

Price

US dollars per million tokens · most-agreed

Input
$0.371
Output
$1.484
Cache read
$0.1275
Cache write
$0.30

Available from 17 hosts

HostIn $/MOut $/MCache rdCache wrContextOutputRetires
llamafreefree128K4K
lambda_ai$0.05$0.10131K8K
abacus · meta-llama$0.14$0.591M8K
io-net · meta-llama$0.15$0.60$0.075$0.30430K4K
deepinfra$0.20$0.801M1M
deepinfra · meta-llama$0.20$0.801M16K
azure$0.25$1.001M16K
azure-cognitive-services$0.25$1.001M16K
hyper$0.255$0.8365$0.1275430K43K
novita$0.27$0.851M8K
novita-ai · meta-llama$0.27$0.851M8K
together_ai$0.27$0.852026-03-31
watsonx$0.371$1.484131K8K
watsonx · meta-llama$0.371$1.484131K8K
oci · oci$0.72$0.721M8K
azure_ai$1.41$0.351M16K
meta_llama · meta-llama1M4K

Price history

append-only observations · a listing writes a row only when its price moves

DateHostField$/M
2026-09-02hyperInput$0.274
2026-09-12hyperInput$0.255
2026-09-02hyperOutput$0.8992
2026-09-12hyperOutput$0.8365
2026-09-11hyperCache read$0.137
2026-09-12hyperCache read$0.1275

More from Meta

most-hosted first