modelbenchmark.io

Llama 3.1 8b Instruct

Meta · llama-3-1-8b-instruct

Compare

Compact open Llama model for lightweight chat, drafting, and self-hosting

Specification

most-agreed values

Context
131K
Max output
16K
Released
2024-07-23
Knowledge cutoff
2023-12
Retires
Open weights
yes
Input
text
Output
text

Price

US dollars per million tokens · most-agreed

Input
$0.0544
Output
$0.085
Cache read
$0.0272
Cache write
$0.20

Available from 19 hosts

HostIn $/MOut $/MCache rdCache wrContextOutputRetires
nvidiafreefree16K4K
helicone$0.02$0.0516K16K
kilo · meta-llama$0.02$0.04131K118K
novita$0.02$0.0516K16K
novita-ai · meta-llama$0.02$0.0516K16K
inference$0.025$0.02516K4K
nscale$0.03$0.03
openrouter$0.05$0.08$0.025131K118K
openrouter · meta-llama$0.05$0.08$0.025131K118K
nano-gpt · meta-llama$0.0544$0.085$0.0272131K16K
huggingface · meta-llama$0.06$0.06131K4K
ovhcloud · ovhcloud$0.10$0.10131K131K
cortecs$0.167$0.167128K128K
perplexity$0.20$0.20131K131K
pioneer · meta-llama$0.20$0.20$0.20$0.20128K8K
merge-gateway$0.22$0.22128K2K
wandb · meta-llama$0.22$0.22$0.22131K131K
wandb · wandb$0.22$0.22128K128K
oci · oci$0.72$0.72128K4K

Other listings of this model

same model, different key

More from Meta

most-hosted first