modelbenchmark.io

Nemotron 3 Ultra

NVIDIA · nemotron-3-ultra-550b

Compare

Flagship Nemotron model for high-throughput reasoning and complex agents

Specification

most-agreed values

Context
131K
Max output
131K
Released
2026-06-04
Knowledge cutoff
Retires
Open weights
no
Input
text
Output
text

Price

US dollars per million tokens · most-agreed

Input
$0.90
Output
$1.70
Cache read
$0.10

Available from 3 hosts

HostIn $/MOut $/MCache rdCache wrContextOutputRetires
llmgateway$0.50$2.20$0.101M128K
llmgateway-providers$0.50$2.20$0.10262K262K
digitalocean$0.90$1.70131K131K

Other listings of this model

same model, different key

More from NVIDIA

most-hosted first