modelbenchmark.io

Llama 3.1 8B (decentralized)

Meta · meta-llama-3-1-8b-instruct-fp8

Compare

Compact GPT model for low-latency assistance and high-volume workloads

Specification

most-agreed values

Context
128K
Max output
16K
Released
2024-01-01
Knowledge cutoff
Retires
Open weights
yes
Input
text
Output
text

Price

US dollars per million tokens · most-agreed

Input
$0.02
Output
$0.03
Cache read
$0.01

Available from 1 host

HostIn $/MOut $/MCache rdCache wrContextOutputRetires
nano-gpt$0.02$0.03$0.01128K16K

More from Meta

most-hosted first