Nemotron-Lightning-3.5-30B-A3B is a 30B-parameter Mixture-of-Experts language model (3B active) from NVIDIA's Nemotron-H family, built on a hybrid Mamba-Transformer architecture for efficient long-context inference. Like other models in the family, it responds to queries by first generating a reasoning trace and then concluding with a final response, with reasoning behavior configurable through a flag in the chat template. It includes a multi-token prediction (MTP) speculative decoding head for low-latency serving.
Specification
most-agreed values
Context
262K
Max output
262K
Released
2026-08-15
Knowledge cutoff
—
Retires
—
Open weights
no
Input
text
Output
text
Price
US dollars per million tokens · most-agreed
Input
$0.05
Output
$0.20
Cache read
$0.01
Available from 3 hosts
| Host | In $/M | Out $/M | Cache rd | Cache wr | Context | Output | Retires |
|---|---|---|---|---|---|---|---|
| fireworks_ai | $0.05 | $0.20 | $0.01 | — | 262K | 33K | — |
| fireworks-ai · accounts | $0.05 | $0.20 | $0.01 | — | 262K | 262K | — |
| requesty | $0.05 | $0.20 | $0.01 | — | 262K | 262K | — |
More from NVIDIA
most-hosted first