Gemini 3.5 Flash

Google released 2026-05-19 proprietary

$3.38 per million tokens, balanced

Gemini 3.7 Flash is both better and cheaper.

It scores +4 higher and costs 78% less ($0.750/M against $3.38/M) on this workload — and it does everything this model does.

GLM 5.3 is cheaper still (36% less) but drops no image, video, file, audio input.

The same model via batch is 50% cheaper ($1.69/M) — same weights, different latency.

Every price dimension

Input$1.50/M
Output$9.00/M
Cached input$0.150/M
Cache write$0.083/M
Reasoning$9.00/M
Web search$0.014/call
Batch discount50%

Reasoning tokens are billed separately at $9.00/M, on top of output. On a reasoning-heavy workload this can be 40% of the bill and it does not appear in the advertised price.

Independent scores

Intelligence52
Coding70.1
Agentic39.7
LMArena Elo1482.6high

Sources are listed separately rather than averaged. Across the 62 models both have scored they correlate at r = 0.806 — close agreement overall, but four models rank very differently between them, and a blended score would hide exactly those.

Best rankings by task

  • agenticslides(python-pptx) #3of 16 1242
  • pptxslides #3of 14 1244
  • agenticslides #4of 16 1244
  • svg #5of 98 1292
  • agentichtmlslides #7of 16 1162
  • agenticslides(html) #7of 16 1162
  • asciiart #8of 83 1285
  • python-pptxslides #9of 35 1247

Capability

Context window1.0M
Max output66K
Input modestext, image, video, file, audio
Tool useyes
Reasoningalways on
Knowledge cutoff2025-01-01
Open weightsno

Provenance

Price sourceopenrouter.ai
Fetched2026-08-24
Quality dataverified
Cross-checkedvendor page

Gemini 3.5 Flash is Google's high-efficiency multimodal model, bringing near-Pro level coding and reasoning at Flash-tier cost and speed. It is highly optimized for coding proficiency and parallel agentic execution...

Markdown for LLMs