---
title: "Google: Gemini 3.1 Pro Preview — price, capability and what beats it · Undominated.ai"
canonical: https://undominated.ai/models/google__gemini-3.1-pro-preview/
description: "Google: Gemini 3.1 Pro Preview: $2.00/M in, $12.00/M out. Independent capability scores, context-tier pricing, and the models that are both better and cheaper."
---

# Google: Gemini 3.1 Pro Preview — price, capability and what beats it · Undominated.ai

> Google: Gemini 3.1 Pro Preview: $2.00/M in, $12.00/M out. Independent capability scores, context-tier pricing, and the models that are both better and cheaper.

[Leaderboard](/) / Google

# Gemini 3.1 Pro Preview

Google · released 2026-02-19 · proprietary

 $4.50 per million tokens, balanced Balanced Summarise Chat Code gen Agentic

## Gemini 3.7 Flash is both better and cheaper.

It scores **+8.3** higher and costs **83% less** ($0.750/M against $4.50/M) on this workload — and it does everything this model does.

[GLM 5.3](/models/z-ai__glm-5.3) is cheaper still (52% less) but drops no audio, file, image, video input.

The same model via **batch** is 50% cheaper ($2.25/M) — same weights, different latency.

## Our take

 editorial — not a measurement

Google's frontier offering at $2/$12 under 200k tokens and $4/$18 above it, with an LMArena score of 1486. Unlike OpenAI, Google publishes the threshold, which makes it budgetable. The catch nobody reads: Gemini's context caching is not just a cheaper read rate, it carries an hourly storage fee — $4.50 per MTok per hour on this model — so a cache you hold open across a workday can cost more than the tokens it saves. Model the storage term explicitly before enabling caching. Genuinely strong multimodal support (audio, video, image, file, text).

### Strengths

 - LMArena Arena Score 1486
- Published 200k tier threshold — budgetable, unlike OpenAI
- Broadest input modality set here: text, image, audio, video, file
- 1M+ context window and 50% batch discount

### Weaknesses

 - Long-context tier doubles input to $4 and raises output to $18
- Context caching carries a $4.50/MTok/hour storage fee on top of read costs
- Still a preview-labelled model

### Reach for it when

 - Multimodal work involving audio or video
- Long-context tasks where you can stay under 200k
- Google Cloud-native stacks

### Avoid it if

 - You want to hold large caches open for hours
- You need a GA rather than preview model

Sources: ai.google.dev · arena.ai

### Every price dimension

| Input | $2.00 /M |
| --- | --- |
| Output | $12.00 /M |
| Cached input | $0.200 /M |
| Cache write | $0.375 /M |
| Reasoning | $12.00 /M |
| Web search | $0.014 /call |
| Batch discount | 50% |

**Past 200,000 tokens the price changes.** Input goes to $4.00/M (2×) and output to $18.00/M. The headline rate does not apply to a long-context workload.

**Reasoning tokens are billed separately** at $12.00/M, on top of output. On a reasoning-heavy workload this can be 40% of the bill and it does not appear in the advertised price.

### Independent scores

| Intelligence | 47.7 |
| --- | --- |
| Coding | 68.8 |
| Agentic | 23 |

Sources are listed separately rather than averaged. Across the 62 models both have scored they correlate at r = 0.806 — close agreement overall, but four models rank very differently between them, and a blended score would hide exactly those.

#### Best rankings by task

 - svg #4 of 98 1328
- asciiart #5 of 83 1299
- agentichtmlslides #5 of 16 1226
- agenticslides(html) #5 of 16 1219
- godotgamedev #6 of 43 1236
- agenticslides #8 of 16 1112
- agenticslides(python-pptx) #8 of 16 1107
- pptxslides #8 of 14 1110

### Capability

| Context window | 1.0M |
| --- | --- |
| Max output | 66K |
| Input modes | audio, file, image, text, video |
| Tool use | yes |
| Reasoning | always on |
| Open weights | no |

### Provenance

| Price source | openrouter.ai |
| --- | --- |
| Fetched | 2026-08-24 |
| Quality data | verified |
| Cross-checked | vendor page |

Gemini 3.1 Pro Preview is Google’s frontier reasoning model, delivering enhanced software engineering performance, improved agentic reliability, and more efficient token usage across complex workflows. Building on the multimodal foundation...
