Nemotron 3 Ultra

NVIDIA released 2026-06-04 nvidia-open-model

$1.35 per million tokens, balanced

DeepSeek V4 Flash 0731 is both better and cheaper.

It scores +13.5 higher and costs 92% less ($0.105/M against $1.35/M) on this workload — and it does everything this model does.

Hy3 preview is cheaper still (79% less) but drops 512K → 262K context.

The same model via free is 100% cheaper (free/M) — same weights, different latency.

Every price dimension

Input$0.600/M
Output$3.60/M
Cached input$0.200/M

Independent scores

Intelligence38.3
Coding49.3
Agentic27.5

Sources are listed separately rather than averaged. Across the 62 models both have scored they correlate at r = 0.806 — close agreement overall, but four models rank very differently between them, and a blended score would hide exactly those.

Capability

Context window512K
Input modestext
Tool useyes
Reasoningoptional
Open weightsyes

Provenance

Price sourceopenrouter.ai
Fetched2026-08-24
Quality dataverified
Cross-checkedvendor page

NVIDIA Nemotron 3 Ultra is an open frontier-reasoning and orchestration model from NVIDIA, with 55B active parameters out of 550B total (MoE). Built on a hybrid Transformer-Mamba mixture-of-experts architecture, it...

Markdown for LLMs