modelbenchmark.io

nvidia-nemotron-3-super-120b-a12b

NVIDIA · nvidia-nemotron-3-super-120b-a12b

Compare

NVIDIA Nemotron 3 Super is a hybrid Mixture-of-Experts (MoE) model engineered for highest compute efficiency and accuracy in multi-agent applications and specialized agentic systems. It is optimized to run many collaborating agents per application on a single GPU, delivering high accuracy for reasoning, tool use, and instruction following.

Specification

most-agreed values

Context
262K
Max output
262K
Released
2026-03-11
Knowledge cutoff
Retires
Open weights
no
Input
text
Output
text

Price

US dollars per million tokens · most-agreed

Input
$0.10
Output
$0.50
Cache read
$0.15

Available from 3 hosts

HostIn $/MOut $/MCache rdCache wrContextOutputRetires
deepinfra$0.085$0.40262K
requesty$0.10$0.50262K262K
crusoe$0.30$2.40$0.15262K262K

Other listings of this model

this page is one of them

The corpus files nvidia-nemotron-3-super-120b-a12b under several keys. This page is the nvidia-nemotron-3-super-120b-a12b listing. The full record — every host, every price and every benchmark score — is on the main nvidia-nemotron-3-super-120b-a12b page.

More from NVIDIA

most-hosted first