Gemini 3.1 Pro Preview
Google released 2026-02-19 proprietary
Gemini 3.7 Flash is both better and cheaper.
It scores +8.3 higher and costs 83% less ($0.750/M against $4.50/M) on this workload — and it does everything this model does.
GLM 5.3 is cheaper still (52% less) but drops no audio, file, image, video input.
The same model via batch is 50% cheaper ($2.25/M) — same weights, different latency.
Our take
editorial — not a measurementGoogle's frontier offering at $2/$12 under 200k tokens and $4/$18 above it, with an LMArena score of 1486. Unlike OpenAI, Google publishes the threshold, which makes it budgetable. The catch nobody reads: Gemini's context caching is not just a cheaper read rate, it carries an hourly storage fee — $4.50 per MTok per hour on this model — so a cache you hold open across a workday can cost more than the tokens it saves. Model the storage term explicitly before enabling caching. Genuinely strong multimodal support (audio, video, image, file, text).
Strengths
- LMArena Arena Score 1486
- Published 200k tier threshold — budgetable, unlike OpenAI
- Broadest input modality set here: text, image, audio, video, file
- 1M+ context window and 50% batch discount
Weaknesses
- Long-context tier doubles input to $4 and raises output to $18
- Context caching carries a $4.50/MTok/hour storage fee on top of read costs
- Still a preview-labelled model
Reach for it when
- Multimodal work involving audio or video
- Long-context tasks where you can stay under 200k
- Google Cloud-native stacks
Avoid it if
- You want to hold large caches open for hours
- You need a GA rather than preview model
Sources: ai.google.devarena.ai
Every price dimension
| Input | $2.00/M |
|---|---|
| Output | $12.00/M |
| Cached input | $0.200/M |
| Cache write | $0.375/M |
| Reasoning | $12.00/M |
| Web search | $0.014/call |
| Batch discount | 50% |
Past 200,000 tokens the price changes. Input goes to $4.00/M (2×) and output to $18.00/M. The headline rate does not apply to a long-context workload.
Reasoning tokens are billed separately at $12.00/M, on top of output. On a reasoning-heavy workload this can be 40% of the bill and it does not appear in the advertised price.
Independent scores
| Intelligence | 47.7 |
|---|---|
| Coding | 68.8 |
| Agentic | 23 |
Sources are listed separately rather than averaged. Across the 62 models both have scored they correlate at r = 0.806 — close agreement overall, but four models rank very differently between them, and a blended score would hide exactly those.
Best rankings by task
- svg #4of 98 1328
- asciiart #5of 83 1299
- agentichtmlslides #5of 16 1226
- agenticslides(html) #5of 16 1219
- godotgamedev #6of 43 1236
- agenticslides #8of 16 1112
- agenticslides(python-pptx) #8of 16 1107
- pptxslides #8of 14 1110
Capability
| Context window | 1.0M |
|---|---|
| Max output | 66K |
| Input modes | audio, file, image, text, video |
| Tool use | yes |
| Reasoning | always on |
| Open weights | no |
Provenance
| Price source | openrouter.ai |
|---|---|
| Fetched | 2026-08-24 |
| Quality data | verified |
| Cross-checked | vendor page |
Gemini 3.1 Pro Preview is Google’s frontier reasoning model, delivering enhanced software engineering performance, improved agentic reliability, and more efficient token usage across complex workflows. Building on the multimodal foundation...