Open Llama instruction model for multilingual chat, reasoning, and coding
Specification
most-agreed values
Context
128K
Max output
4K
Released
2024-07-23
Knowledge cutoff
2023-12-31
Retires
2026-06-13
Open weights
yes
Input
text
Output
text
Price
US dollars per million tokens · most-agreed
Input
$0.02
Output
$0.05
Cache read
$0.15
Cache write
$0.15
Available from 8 hosts
| Host | In $/M | Out $/M | Cache rd | Cache wr | Context | Output | Retires |
|---|---|---|---|---|---|---|---|
| abacus · meta-llama | $0.02 | $0.05 | — | — | 128K | 4K | — |
| nebius | $0.02 | $0.06 | — | — | 128K | 128K | — |
| deepinfra | $0.03 | $0.05 | — | — | 131K | 131K | — |
| sambanova · sambanova | $0.10 | $0.20 | — | — | 16K | 16K | 2026-04-14 |
| hyperbolic | $0.12 | $0.30 | — | — | 33K | 33K | — |
| databricks | $0.15 | $0.45 | $0.15 | $0.15 | 200K | 128K | — |
| neon | $0.15 | $0.45 | — | — | 131K | 8K | — |
| azure_ai | $0.30 | $0.61 | — | — | 128K | 2K | 2026-06-13 |
Other listings of this model
this page is one of them
The corpus files Llama 3.1 8B Instruct under several keys. This page is the meta-llama-3-1-8b-instruct listing. The full record — every host, every price and every benchmark score — is on the main Llama 3.1 8B Instruct page.
More from Meta
most-hosted first