Gemini 3.1 Pro Preview

Google released 2026-02-19 proprietary

$4.50 per million tokens, balanced

Gemini 3.7 Flash is both better and cheaper.

It scores +8.3 higher and costs 83% less ($0.750/M against $4.50/M) on this workload — and it does everything this model does.

GLM 5.3 is cheaper still (52% less) but drops no audio, file, image, video input.

The same model via batch is 50% cheaper ($2.25/M) — same weights, different latency.

Our take

editorial — not a measurement

Google's frontier offering at $2/$12 under 200k tokens and $4/$18 above it, with an LMArena score of 1486. Unlike OpenAI, Google publishes the threshold, which makes it budgetable. The catch nobody reads: Gemini's context caching is not just a cheaper read rate, it carries an hourly storage fee — $4.50 per MTok per hour on this model — so a cache you hold open across a workday can cost more than the tokens it saves. Model the storage term explicitly before enabling caching. Genuinely strong multimodal support (audio, video, image, file, text).

Strengths

  • LMArena Arena Score 1486
  • Published 200k tier threshold — budgetable, unlike OpenAI
  • Broadest input modality set here: text, image, audio, video, file
  • 1M+ context window and 50% batch discount

Weaknesses

  • Long-context tier doubles input to $4 and raises output to $18
  • Context caching carries a $4.50/MTok/hour storage fee on top of read costs
  • Still a preview-labelled model

Reach for it when

  • Multimodal work involving audio or video
  • Long-context tasks where you can stay under 200k
  • Google Cloud-native stacks

Avoid it if

  • You want to hold large caches open for hours
  • You need a GA rather than preview model

Sources: ai.google.devarena.ai

Every price dimension

Input$2.00/M
Output$12.00/M
Cached input$0.200/M
Cache write$0.375/M
Reasoning$12.00/M
Web search$0.014/call
Batch discount50%

Past 200,000 tokens the price changes. Input goes to $4.00/M (2×) and output to $18.00/M. The headline rate does not apply to a long-context workload.

Reasoning tokens are billed separately at $12.00/M, on top of output. On a reasoning-heavy workload this can be 40% of the bill and it does not appear in the advertised price.

Independent scores

Intelligence47.7
Coding68.8
Agentic23

Sources are listed separately rather than averaged. Across the 62 models both have scored they correlate at r = 0.806 — close agreement overall, but four models rank very differently between them, and a blended score would hide exactly those.

Best rankings by task

  • svg #4of 98 1328
  • asciiart #5of 83 1299
  • agentichtmlslides #5of 16 1226
  • agenticslides(html) #5of 16 1219
  • godotgamedev #6of 43 1236
  • agenticslides #8of 16 1112
  • agenticslides(python-pptx) #8of 16 1107
  • pptxslides #8of 14 1110

Capability

Context window1.0M
Max output66K
Input modesaudio, file, image, text, video
Tool useyes
Reasoningalways on
Open weightsno

Provenance

Price sourceopenrouter.ai
Fetched2026-08-24
Quality dataverified
Cross-checkedvendor page

Gemini 3.1 Pro Preview is Google’s frontier reasoning model, delivering enhanced software engineering performance, improved agentic reliability, and more efficient token usage across complex workflows. Building on the multimodal foundation...

Markdown for LLMs