Gemini 3.7 Flash
Google released 2026-08-13 proprietary
Nothing is both better and cheaper.
Under a balanced workload, no other model in the catalogue scores higher and costs less. This model is on the value frontier.
The same model via batch is 50% cheaper ($0.375/M) — same weights, different latency.
Our take
editorial — not a measurementThe best-value Google model right now, but only until 31 December 2026. It is $0.75/$3.75 on promotional pricing that Google has explicitly dated: on 1 January 2027 it becomes $1.50/$7.50, a 100% increase already on the published schedule. Anyone building a cost model on today's rate needs to plan for the doubling. LMArena 1490 at high effort is remarkable for a Flash-tier model — that beats Gemini 3.1 Pro's 1486. Note OpenRouter lists this at $0.375/$1.875, which is Google's batch rate, so aggregator data disagrees with the vendor here.
Strengths
- LMArena 1490 at high effort — outscores Gemini 3.1 Pro
- $0.75/$3.75 promotional, cheap for the capability
- Full multimodal input including audio and video
- 1M+ context, 50% batch discount
Weaknesses
- Price doubles to $1.50/$7.50 on 2027-01-01 — already scheduled
- OpenRouter's listed rate conflicts with Google's own page
- Context caching storage fees apply as with all Gemini models
Reach for it when
- High-volume multimodal work through 2026
- Cheap access to near-Pro quality
Avoid it if
- You are committing to a multi-year cost model
- You need pricing stability past 2026
Sources: ai.google.devarena.ai
Every price dimension
| Input | $0.375/M |
|---|---|
| Output | $1.88/M |
| Cached input | $0.037/M |
| Cache write | $0.021/M |
| Reasoning | $1.88/M |
| Web search | $0.014/call |
| Batch discount | 50% |
Reasoning tokens are billed separately at $1.88/M, on top of output. On a reasoning-heavy workload this can be 40% of the bill and it does not appear in the advertised price.
Independent scores
| Intelligence | 56 |
|---|---|
| Coding | 76.1 |
| Agentic | 45.1 |
| LMArena Elo | 1490.2high |
Sources are listed separately rather than averaged. Across the 62 models both have scored they correlate at r = 0.806 — close agreement overall, but four models rank very differently between them, and a blended score would hide exactly those.
Best rankings by task
- agenticgamedev #2of 34 1277
- androidnative #3of 54 1277
- codecategories #5of 145 1332
- gamedev #5of 144 1348
- dataviz #5of 142 1351
- 3d #5of 136 1357
- website #6of 151 1323
- mobileapps #10of 56 1247
Capability
| Context window | 1.0M |
|---|---|
| Max output | 66K |
| Input modes | text, image, video, file, audio |
| Tool use | yes |
| Reasoning | always on |
| Open weights | no |
Provenance
| Price source | openrouter.ai |
|---|---|
| Fetched | 2026-08-24 |
| Quality data | verified |
| Cross-checked | vendor page |
Gemini 3.7 Flash is a multimodal model from Google for fast agentic workflows, coding, and complex multi-step reasoning. It is designed for tasks that require responsive performance and reliable multi-step...