modelbenchmark.io

nemotron-3-ultra

NVIDIA · nemotron-3-ultra

Compare

Largest Nemotron 3 model for maximum open-weight reasoning and agent accuracy

Specification

most-agreed values

Context
262K
Max output
128K
Released
2026-06-04
Knowledge cutoff
Retires
Open weights
yes
Input
text
Output
text

Price

US dollars per million tokens · most-agreed

Input
$0.10
Output
$3.00
Cache read
$0.10

Quality

4 benchmarks · 1 source

Composite
75th
percentile of 334 scored models
Rank
83 / 334
±0.10 sd
Evidence
4 × 1
one source, or fewer than 3 benchmarks
Effort range
one configuration only
reasoning71st3/4 bench
math91st1/7 bench

Every score

one row per source and configuration — nothing averaged away

BenchmarkScoreConfigurationSourceRun
OTIS Mock AIME 2024-202586.7 ±5.1defaultEpoch AI2026-08-10
GPQA diamond85.3 ±2.5defaultEpoch AI2026-08-10
Mystery Game Puzzles20.0 ±4.0defaultEpoch AI2026-08-27
Chess Puzzles12.0 ±3.3defaultEpoch AI2026-08-10

The leading figure for a benchmark is the median across its configurations, so one heroic high-effort run cannot set the number.

Available from 2 hosts

HostIn $/MOut $/MCache rdCache wrContextOutputRetires
ollama-cloud$0.10$3.00$0.10262K128K
routing-run$0.10$0.10131K32K

Other listings of this model

this page is one of them

The corpus files nemotron-3-ultra under several keys. This page is the nemotron-3-ultra listing. The full record — every host, every price and every benchmark score — is on the main nemotron-3-ultra page.

More from NVIDIA

most-hosted first