NVIDIA Nemotron 3.5 Lightning 30B-A3B is a hybrid Mamba-2 + MoE + Attention model with 30B total and 3B active parameters, pre-trained on over 20T tokens with an NVFP4 recipe and Multi-Token Prediction for fast generation. Up to 1M token context for long-running autonomous agents, sub-agent workhorse deployments, and agentic workflows. Supports reasoning and tool calling. English and coding languages plus Spanish, French, German, Italian, and Japanese. Open weights under the OpenMDW License Agreement v1.1. Part of the NVIDIA Nemotron family.
Specification
most-agreed values
Context
1M
Max output
66K
Released
2026-08-11
Knowledge cutoff
—
Retires
—
Open weights
no
Input
text
Output
text
Price
US dollars per million tokens · most-agreed
Input
free
Output
free
Cache read
$0.0011
Available from 4 hosts
| Host | In $/M | Out $/M | Cache rd | Cache wr | Context | Output | Retires |
|---|---|---|---|---|---|---|---|
| merge-gateway | free | free | — | — | 1M | 262K | — |
| nvidia | free | free | — | — | 262K | 262K | — |
| requesty | free | free | — | — | 1M | 66K | — |
| perplexity | $0.0115 | $0.17 | $0.0011 | — | — | — | — |
Other listings of this model
same model, different key
More from NVIDIA
most-hosted first