modelbenchmark.io

Nemotron 3 Ultra 550B A55B

NVIDIA · nvidia-nemotron-3-ultra-550b-a55b-bf16

Compare

Largest Nemotron 3 model for maximum open-weight reasoning and agent accuracy

Specification

most-agreed values

Context
1M
Max output
65K
Released
2026-06-04
Knowledge cutoff
Retires
Open weights
yes
Input
text
Output
text

Price

US dollars per million tokens · most-agreed

Input
$0.50
Output
$2.50
Cache read
$0.15
Cache write
$0.50

Available from 1 host

HostIn $/MOut $/MCache rdCache wrContextOutputRetires
pioneer$0.50$2.50$0.15$0.501M65K

More from NVIDIA

most-hosted first