Gemini 3.5 Flash Lite
Google released 2026-07-21 proprietary
Gemini 3.7 Flash is both better and cheaper.
It scores +18.6 higher and costs 12% less ($0.750/M against $0.850/M) on this workload — and it does everything this model does.
DeepSeek V4 Flash 0731 is cheaper still (88% less) but drops no image, video, file, audio input.
The same model via batch is 50% cheaper ($0.425/M) — same weights, different latency.
Every price dimension
| Input | $0.300/M |
|---|---|
| Output | $2.50/M |
| Cached input | $0.030/M |
| Cache write | $0.083/M |
| Reasoning | $2.50/M |
| Web search | $0.014/call |
| Batch discount | 50% |
Reasoning tokens are billed separately at $2.50/M, on top of output. On a reasoning-heavy workload this can be 40% of the bill and it does not appear in the advertised price.
Independent scores
| Intelligence | 37.4 |
|---|---|
| Coding | 49.3 |
| Agentic | 27.2 |
| LMArena Elo | 1436.5default |
Sources are listed separately rather than averaged. Across the 62 models both have scored they correlate at r = 0.806 — close agreement overall, but four models rank very differently between them, and a blended score would hide exactly those.
Capability
| Context window | 1.0M |
|---|---|
| Max output | 66K |
| Input modes | text, image, video, file, audio |
| Tool use | yes |
| Reasoning | always on |
| Open weights | no |
Provenance
| Price source | openrouter.ai |
|---|---|
| Fetched | 2026-08-24 |
| Quality data | verified |
| Cross-checked | vendor page |
Gemini 3.5 Flash Lite is a high-efficiency model from Google with upgraded agentic capabilities. It is suited for subagents that execute focused tasks within complex, multi-agent workflows.