modelbenchmark.io

Nemotron Ultra

NVIDIA · nvidia-nemotron-3-ultra-550b-a55b

Compare

Largest Nemotron 3 model for maximum open-weight reasoning and agent accuracy

Specification

most-agreed values

Context
203K
Max output
203K
Released
2026-06-04
Knowledge cutoff
Retires
Open weights
yes
Input
text
Output
text

Price

US dollars per million tokens · most-agreed

Input
$0.60
Output
$2.40
Cache read
$0.12

Available from 5 hosts

HostIn $/MOut $/MCache rdCache wrContextOutputRetires
deepinfra$0.50$2.20$0.10262K
wandb$0.50$2.15$0.10262K262K
baseten$0.60$2.40$0.12203K203K
venice$0.625$3.125$0.1875256K33K
wandb · wandb$0.75$2.75$0.15262K

Price history

append-only observations · a listing writes a row only when its price moves

DateHostField$/M
2026-09-02wandbCache read$0.15
2026-09-15wandbCache read$0.10
2026-09-02wandbInput$0.75
2026-09-15wandbInput$0.50
2026-09-02wandbOutput$2.75
2026-09-15wandbOutput$2.15

More from NVIDIA

most-hosted first