Nemotron 3.5 Lightning

NVIDIA released 2026-08-11

$0.110 per million tokens, balanced

DeepSeek V4 Flash 0731 is both better and cheaper.

It scores +28.2 higher and costs 5% less ($0.105/M against $0.110/M) on this workload — and it does everything this model does.

Ling-3.0-flash is cheaper still (71% less) but drops 131K → 33K max output.

The same model via free is 100% cheaper (free/M) — same weights, different latency.

Every price dimension

Input$0.080/M
Output$0.200/M
Cached input$0.040/M

Independent scores

Intelligence23.6
Coding26.8
Agentic13.8

Sources are listed separately rather than averaged. Across the 62 models both have scored they correlate at r = 0.806 — close agreement overall, but four models rank very differently between them, and a blended score would hide exactly those.

Capability

Context window262K
Max output131K
Input modestext
Tool useyes
Reasoningoptional

Provenance

Price sourceopenrouter.ai
Fetched2026-08-24
Quality dataverified

NVIDIA Nemotron 3.5 Lightning is an open mixture-of-experts model from NVIDIA, with 3B active parameters out of 30B total. It is suited for high-throughput agentic workloads and specialized tasks that...

Markdown for LLMs