NVIDIA Nemotron 3 Super is a hybrid Mixture-of-Experts (MoE) model engineered for highest compute efficiency and accuracy in multi-agent applications and specialized agentic systems. It is optimized to run many collaborating agents per application on a single GPU, delivering high accuracy for reasoning, tool use, and instruction following.
Specification
most-agreed values
Context
262K
Max output
262K
Released
2026-03-11
Knowledge cutoff
—
Retires
—
Open weights
no
Input
text
Output
text
Price
US dollars per million tokens · most-agreed
Input
$0.10
Output
$0.50
Cache read
$0.15
Available from 3 hosts
Other listings of this model
this page is one of them
The corpus files nvidia-nemotron-3-super-120b-a12b under several keys. This page is the nvidia-nemotron-3-super-120b-a12b listing. The full record — every host, every price and every benchmark score — is on the main nvidia-nemotron-3-super-120b-a12b page.
More from NVIDIA
most-hosted first