Gemini 3.5 Flash Lite

Google released 2026-07-21 proprietary

$0.850 per million tokens, balanced

Gemini 3.7 Flash is both better and cheaper.

It scores +18.6 higher and costs 12% less ($0.750/M against $0.850/M) on this workload — and it does everything this model does.

DeepSeek V4 Flash 0731 is cheaper still (88% less) but drops no image, video, file, audio input.

The same model via batch is 50% cheaper ($0.425/M) — same weights, different latency.

Every price dimension

Input$0.300/M
Output$2.50/M
Cached input$0.030/M
Cache write$0.083/M
Reasoning$2.50/M
Web search$0.014/call
Batch discount50%

Reasoning tokens are billed separately at $2.50/M, on top of output. On a reasoning-heavy workload this can be 40% of the bill and it does not appear in the advertised price.

Independent scores

Intelligence37.4
Coding49.3
Agentic27.2
LMArena Elo1436.5default

Sources are listed separately rather than averaged. Across the 62 models both have scored they correlate at r = 0.806 — close agreement overall, but four models rank very differently between them, and a blended score would hide exactly those.

Capability

Context window1.0M
Max output66K
Input modestext, image, video, file, audio
Tool useyes
Reasoningalways on
Open weightsno

Provenance

Price sourceopenrouter.ai
Fetched2026-08-24
Quality dataverified
Cross-checkedvendor page

Gemini 3.5 Flash Lite is a high-efficiency model from Google with upgraded agentic capabilities. It is suited for subagents that execute focused tasks within complex, multi-agent workflows.

Markdown for LLMs