modelbenchmark.io

Llama 3.3 70B Instruct fp8 Fast

Meta · llama-3-3-70b-instruct-fp8-fast

Compare

Popular open Llama workhorse for multilingual chat, coding, and self-hosting

Specification

most-agreed values

Context
24K
Max output
24K
Released
2024-12-06
Knowledge cutoff
2023-12
Retires
Open weights
yes
Input
text
Output
text

Price

US dollars per million tokens · most-agreed

Input
$0.293
Output
$2.253

Rate limits

Requests per minute
300

Available from 2 hosts

HostIn $/MOut $/MCache rdCache wrContextOutputRetires
cloudflare$0.293$2.25324K24K
cloudflare-workers-ai · cf$0.293$2.25324K24K

More from Meta

most-hosted first