Gemini 3.5 Flash
Google released 2026-05-19 proprietary
Gemini 3.7 Flash is both better and cheaper.
It scores +4 higher and costs 78% less ($0.750/M against $3.38/M) on this workload — and it does everything this model does.
GLM 5.3 is cheaper still (36% less) but drops no image, video, file, audio input.
The same model via batch is 50% cheaper ($1.69/M) — same weights, different latency.
Every price dimension
| Input | $1.50/M |
|---|---|
| Output | $9.00/M |
| Cached input | $0.150/M |
| Cache write | $0.083/M |
| Reasoning | $9.00/M |
| Web search | $0.014/call |
| Batch discount | 50% |
Reasoning tokens are billed separately at $9.00/M, on top of output. On a reasoning-heavy workload this can be 40% of the bill and it does not appear in the advertised price.
Independent scores
| Intelligence | 52 |
|---|---|
| Coding | 70.1 |
| Agentic | 39.7 |
| LMArena Elo | 1482.6high |
Sources are listed separately rather than averaged. Across the 62 models both have scored they correlate at r = 0.806 — close agreement overall, but four models rank very differently between them, and a blended score would hide exactly those.
Best rankings by task
- agenticslides(python-pptx) #3of 16 1242
- pptxslides #3of 14 1244
- agenticslides #4of 16 1244
- svg #5of 98 1292
- agentichtmlslides #7of 16 1162
- agenticslides(html) #7of 16 1162
- asciiart #8of 83 1285
- python-pptxslides #9of 35 1247
Capability
| Context window | 1.0M |
|---|---|
| Max output | 66K |
| Input modes | text, image, video, file, audio |
| Tool use | yes |
| Reasoning | always on |
| Knowledge cutoff | 2025-01-01 |
| Open weights | no |
Provenance
| Price source | openrouter.ai |
|---|---|
| Fetched | 2026-08-24 |
| Quality data | verified |
| Cross-checked | vendor page |
Gemini 3.5 Flash is Google's high-efficiency multimodal model, bringing near-Pro level coding and reasoning at Flash-tier cost and speed. It is highly optimized for coding proficiency and parallel agentic execution...