modelbenchmark.io

Llama-3.3-70B-Instruct

NVIDIA · llama-3-3-70b-instruct-fp8

Compare

Popular open Llama workhorse for multilingual chat, coding, and self-hosting

Specification

most-agreed values

Context
128K
Max output
4K
Released
2024-12-06
Knowledge cutoff
2023-12
Retires
Open weights
yes
Input
text
Output
text

Price

US dollars per million tokens · most-agreed

Input
$1.15
Output
$1.15

Available from 1 host

HostIn $/MOut $/MCache rdCache wrContextOutputRetires
evroc$1.15$1.15128K4K

Other listings of this model

this page is one of them

The corpus files Llama-3.3-70B-Instruct under several keys. This page is the llama-3-3-70b-instruct-fp8 listing. The full record — every host, every price and every benchmark score — is on the main Llama-3.3-70B-Instruct page.

More from NVIDIA

most-hosted first