modelbenchmark.io

nemotron-lightning-3.5-30b-a3b

NVIDIA · nemotron-lightning-3-5-30b-a3b

Compare

Nemotron-Lightning-3.5-30B-A3B is a 30B-parameter Mixture-of-Experts language model (3B active) from NVIDIA's Nemotron-H family, built on a hybrid Mamba-Transformer architecture for efficient long-context inference. Like other models in the family, it responds to queries by first generating a reasoning trace and then concluding with a final response, with reasoning behavior configurable through a flag in the chat template. It includes a multi-token prediction (MTP) speculative decoding head for low-latency serving.

Specification

most-agreed values

Context
262K
Max output
262K
Released
2026-08-15
Knowledge cutoff
Retires
Open weights
no
Input
text
Output
text

Price

US dollars per million tokens · most-agreed

Input
$0.05
Output
$0.20
Cache read
$0.01

Available from 3 hosts

HostIn $/MOut $/MCache rdCache wrContextOutputRetires
fireworks_ai$0.05$0.20$0.01262K33K
fireworks-ai · accounts$0.05$0.20$0.01262K262K
requesty$0.05$0.20$0.01262K262K

More from NVIDIA

most-hosted first