modelbenchmark.io

Llama 3.1 Nemotron Ultra 253B

NVIDIA · llama-3-1-nemotron-ultra-253b

Compare

Flagship Nemotron model for high-throughput reasoning and complex agents

Specification

most-agreed values

Context
128K
Max output
16K
Released
2025-04-07
Knowledge cutoff
Retires
Open weights
yes
Input
text
Output
text

Price

US dollars per million tokens · most-agreed

Input
free
Output
free

Available from 2 hosts

HostIn $/MOut $/MCache rdCache wrContextOutputRetires
nvidiafreefree128K16K
nebius$0.60$1.80
128K
131K
128K
131K

Price history

append-only observations · a listing writes a row only when its price moves

DateHostField$/M
2026-09-02nebiusInput$0.60
2026-09-03nebiusInput$0.60
2026-09-02nebiusOutput$1.80
2026-09-03nebiusOutput$1.80

More from NVIDIA

most-hosted first