Gemini 3.7 Flash

Google released 2026-08-13 proprietary

$0.750 per million tokens, balanced

Nothing is both better and cheaper.

Under a balanced workload, no other model in the catalogue scores higher and costs less. This model is on the value frontier.

The same model via batch is 50% cheaper ($0.375/M) — same weights, different latency.

Our take

editorial — not a measurement

The best-value Google model right now, but only until 31 December 2026. It is $0.75/$3.75 on promotional pricing that Google has explicitly dated: on 1 January 2027 it becomes $1.50/$7.50, a 100% increase already on the published schedule. Anyone building a cost model on today's rate needs to plan for the doubling. LMArena 1490 at high effort is remarkable for a Flash-tier model — that beats Gemini 3.1 Pro's 1486. Note OpenRouter lists this at $0.375/$1.875, which is Google's batch rate, so aggregator data disagrees with the vendor here.

Strengths

  • LMArena 1490 at high effort — outscores Gemini 3.1 Pro
  • $0.75/$3.75 promotional, cheap for the capability
  • Full multimodal input including audio and video
  • 1M+ context, 50% batch discount

Weaknesses

  • Price doubles to $1.50/$7.50 on 2027-01-01 — already scheduled
  • OpenRouter's listed rate conflicts with Google's own page
  • Context caching storage fees apply as with all Gemini models

Reach for it when

  • High-volume multimodal work through 2026
  • Cheap access to near-Pro quality

Avoid it if

  • You are committing to a multi-year cost model
  • You need pricing stability past 2026

Sources: ai.google.devarena.ai

Every price dimension

Input$0.375/M
Output$1.88/M
Cached input$0.037/M
Cache write$0.021/M
Reasoning$1.88/M
Web search$0.014/call
Batch discount50%

Reasoning tokens are billed separately at $1.88/M, on top of output. On a reasoning-heavy workload this can be 40% of the bill and it does not appear in the advertised price.

Independent scores

Intelligence56
Coding76.1
Agentic45.1
LMArena Elo1490.2high

Sources are listed separately rather than averaged. Across the 62 models both have scored they correlate at r = 0.806 — close agreement overall, but four models rank very differently between them, and a blended score would hide exactly those.

Best rankings by task

  • agenticgamedev #2of 34 1277
  • androidnative #3of 54 1277
  • codecategories #5of 145 1332
  • gamedev #5of 144 1348
  • dataviz #5of 142 1351
  • 3d #5of 136 1357
  • website #6of 151 1323
  • mobileapps #10of 56 1247

Capability

Context window1.0M
Max output66K
Input modestext, image, video, file, audio
Tool useyes
Reasoningalways on
Open weightsno

Provenance

Price sourceopenrouter.ai
Fetched2026-08-24
Quality dataverified
Cross-checkedvendor page

Gemini 3.7 Flash is a multimodal model from Google for fast agentic workflows, coding, and complex multi-step reasoning. It is designed for tasks that require responsive performance and reliable multi-step...

Markdown for LLMs