---
title: "Google: Gemma 4 31B — price, capability and what beats it · Undominated.ai"
canonical: https://undominated.ai/models/google__gemma-4-31b-it/
description: "Google: Gemma 4 31B: $0.140/M in, $0.400/M out. Independent capability scores, and alternatives compared under recorded workload and capability requirements."
---

# Google: Gemma 4 31B — price, capability and what beats it · Undominated.ai

> Google: Gemma 4 31B: $0.140/M in, $0.400/M out. Independent capability scores, and alternatives compared under recorded workload and capability requirements.

[Check base rates for this workload](/check/google__gemma-4-31b-it/) [Cheaper alternatives](/alternatives/google__gemma-4-31b-it/) [Family](/families/gemma/) [Compare models](https://undominated.ai/compare/?models=google%2Fgemma-4-31b-it) Comparison context

Choose up to three standard models in Compare.

Requirements and prompt length apply in Compare. This page keeps its stated price basis and workload controls.

MODEL PROOF

# Gemma 4 31B

Google · 02 Apr 2026 · apache-2.0

 Balanced Summarise Chat Code gen Agentic

BETTER VALUE OPTION

## A stored alternative has an equal-or-higher measured score and an equal-or-lower price.

 **$0.205** / 1M tokens Balanced · 3 tokens in per 1 out

Serving precision differs between offers.

[MiMo-V2.6-Flash](/models/xiaomi__mimo-v2.6-flash/) · Δ score 13 points · 15% lower measured price · $0.175 / 1M tokens

[DeepSeek V4.1 Flash](/models/deepseek__deepseek-v4.1-flash/) is a further stored option with these losses: no video input

Evidence as of 06 Oct 2026

 [openrouter/models](https://openrouter.ai/api/v1/models) [Vendor cross-check](https://openrouter.ai/api/v1/models) [LMArena · cc-by-4.0](https://huggingface.co/datasets/lmarena-ai/leaderboard-dataset) Independent LMArena score 1,443.4 ± 7.4 · 6,133 votes
 Context window 262.14K
 Maximum output 16.38K
 Input modalities image, text, video
 Output modalities text
 Published input price $0.140 / 1M tokens
 Published output price $0.400 / 1M tokens
 Pricing kind fixed

Inspect [complete billing conditions](#billing-details) and endpoint terms below.

Price from Crusoe: the cheapest offer that is serving, at standard delivery and at a declared precision that is not 4-bit (bf16), at its standard rate. The maker’s own precision is not known.

Cheaper right now: $0.153/M at DeepInfra — delivery tier (turbo) · 4-bit It does not pass the like-for-like test, so it never sets a rank.

 Every price dimension · USD per million tokens

| Input | $0.140 /M |
| --- | --- |
| Output | $0.400 /M |
| Cached input | $0.140 /M |
| Cache write | Unknown /M |
| Cache write 1h | Unknown /M |
| Reasoning | Unknown /M |

### Who sells it

The same weights, different shops. Cheapest is not like-for-like when serving precision differs.

Headline cheapest is a lower precision. Like-for-like at the best declared precision is $0.140/M from Crusoe.

 13 provider offers · rates, limits and conditions

Rates are USD per million tokens. Endpoint terms can differ even when the quoted price and precision match. A listed parameter is a provider declaration, not a task-success test.

| Seller | Input | Output | Precision | Uptime (1d) |
| --- | --- | --- | --- | --- |
| Seller DeepInfra deepinfra/turbo Endpoint terms Cached input /M $0.050 Context limit 262,144 tokens Output limit 16,384 tokens Tools Not listed by endpoint Reasoning Listed by endpoint Promotional discount None reported Uptime · last 30 minutes 98.90% | Input $0.090 | Output $0.340 | Precision fp4 | Uptime (1d) 99.29% |
| Seller CoreWeave coreweave/fp4 Endpoint terms Cached input /M $0.100 Context limit 262,144 tokens Output limit 235,929 tokens Tools Listed by endpoint Reasoning Listed by endpoint Promotional discount None reported Uptime · last 30 minutes 99.55% | Input $0.100 | Output $0.340 | Precision fp4 | Uptime (1d) 99.24% |
| Seller Venice venice/fp4 Endpoint terms Cached input /M $0.090 Context limit 256,000 tokens Output limit 8,192 tokens Tools Listed by endpoint Reasoning Listed by endpoint Promotional discount None reported Uptime · last 30 minutes 98.14% | Input $0.120 | Output $0.360 | Precision fp4 | Uptime (1d) 99.41% |
| Seller Chutes chutes/fp4 Endpoint terms Cached input /M $0.012 Context limit 131,072 tokens Output limit 65,536 tokens Tools Listed by endpoint Reasoning Listed by endpoint Promotional discount None reported Uptime · last 30 minutes 95.93% | Input $0.120 | Output $0.370 | Precision fp4 | Uptime (1d) 94.20% |
| Seller Crusoe crusoe/bf16 Endpoint terms Cached input /M $0.140 Context limit 262,144 tokens Output limit 262,141 tokens Tools Listed by endpoint Reasoning Listed by endpoint Promotional discount None reported Uptime · last 30 minutes 99.89% | Input $0.140 | Output $0.400 | Precision bf16 | Uptime (1d) 99.48% |
| Seller Novita novita/bf16 Endpoint terms Cached input /M Unknown Context limit 262,144 tokens Output limit 131,072 tokens Tools Listed by endpoint Reasoning Listed by endpoint Promotional discount None reported Uptime · last 30 minutes 81.87% | Input $0.140 | Output $0.400 | Precision bf16 | Uptime (1d) 82.80% |
| Seller Friendli friendli Endpoint terms Cached input /M Unknown Context limit 262,144 tokens Output limit 8,192 tokens Tools Listed by endpoint Reasoning Listed by endpoint Promotional discount None reported Uptime · last 30 minutes 99.93% | Input $0.140 | Output $0.400 | Precision undeclared | Uptime (1d) 99.59% |
| Seller Parasail parasail/fp8 Endpoint terms Cached input /M $0.060 Context limit 262,144 tokens Output limit 235,929 tokens Tools Listed by endpoint Reasoning Listed by endpoint Promotional discount None reported Uptime · last 30 minutes 98.96% | Input $0.150 | Output $0.400 | Precision fp8 | Uptime (1d) 99.33% |
| Seller DeepInfra deepinfra/fp8 Endpoint terms Cached input /M Unknown Context limit 262,144 tokens Output limit 16,384 tokens Tools Listed by endpoint Reasoning Listed by endpoint Promotional discount None reported Uptime · last 30 minutes 99.52% | Input $0.200 | Output $0.400 | Precision fp8 | Uptime (1d) 98.77% |
| Seller Io Net io-net Endpoint terms Cached input /M $0.181 Context limit 262,144 tokens Output limit 16,384 tokens Tools Listed by endpoint Reasoning Listed by endpoint Promotional discount 5% · already included in these rates Uptime · last 30 minutes 99.35% | Input $0.361 | Output $1.09 | Precision undeclared | Uptime (1d) 99.28% |
| Seller SambaNova sambanova Endpoint terms Cached input /M Unknown Context limit 262,144 tokens Output limit 235,929 tokens Tools Listed by endpoint Reasoning Listed by endpoint Promotional discount None reported Uptime · last 30 minutes 97.95% | Input $0.380 | Output $1.15 | Precision undeclared | Uptime (1d) 98.55% |
| Seller SiliconFlow siliconflow/fp8 Endpoint terms Cached input /M $0.250 Context limit 262,144 tokens Output limit 235,929 tokens Tools Listed by endpoint Reasoning Listed by endpoint Promotional discount None reported Uptime · last 30 minutes 99.18% | Input $0.750 | Output $1.00 | Precision fp8 | Uptime (1d) 76.52% |
| Seller ModelRun modelrun/fp4 Endpoint terms Cached input /M $0.200 Context limit 262,144 tokens Output limit 235,929 tokens Tools Listed by endpoint Reasoning Listed by endpoint Promotional discount None reported Uptime · last 30 minutes 99.36% | Input $0.750 | Output $1.00 | Precision fp4 | Uptime (1d) 98.99% |

### Independent scores

| LMArena Elo | 1443.4 default |
| --- | --- |

LMArena Elo under CC BY 4.0. Absence of another board is not a score of zero.

### Capability

| Context window | 262K |
| --- | --- |
| Max output | 16K |
| Input modes | image, text, video |
| Tool use | yes |
| Reasoning | optional |
| Open weights | yes |

### Provenance

| Price source | [openrouter.ai](https://openrouter.ai/api/v1/models) |
| --- | --- |
| Fetched | 2026-10-06 |
| Quality data | verified |
| Cross-checked | [vendor page](https://openrouter.ai/api/v1/models) |

Gemma 4 31B Instruct is Google DeepMind's 30.7B dense multimodal model supporting text and image input with text output. Features a 256K token context window, configurable thinking/reasoning mode, native function...

## Questions this page answers

 What does Gemma 4 31B cost?

$0.140/M in, $0.400/M out in this catalogue, as of the fetch date on this page. That is the API row, not a subscription.

 Does Gemma 4 31B have an independent quality score?

Gemma 4 31B has an LMArena score in this catalogue.

 What beats Gemma 4 31B?

MiMo-V2.6-Flash has an equal-or-higher measured score and an equal-or-lower price.

 Does this page use Artificial Analysis scores for Gemma 4 31B?

No. Artificial Analysis figures are not published here. Quality on this page is LMArena Elo where a score exists; otherwise the row is unrated.

 Is the cheapest Gemma 4 31B endpoint the same product?

Headline cheapest is a lower precision. Like-for-like at the best declared precision is $0.140/M from Crusoe.

MODEL MONUMENT

## Save or share this proof

 Download proof card · 1200 × 630 Preview and more sharing options

Google: Gemma 4 31B. Balanced workload. Better value option available. Evidence as of 06 Oct 2026.

 Copy model link Download square card · 1080 × 1080 Copy landscape image Share landscape card

## Related

 - Both better and cheaper [MiMo-V2.6-Flash](/models/xiaomi__mimo-v2.6-flash/)
- Higher score and cheaper, with a named trade [DeepSeek V4.1 Flash](/models/deepseek__deepseek-v4.1-flash/)
- Side by side [Gemma 4 26B A4B](/compare/google__gemma-4-26b-a4b-it--vs--google__gemma-4-31b-it/)
- Nearby on the capability axis [DeepSeek V4 Pro 0423](/models/deepseek__deepseek-v4-pro/)
- Same vendor [Gemini 3.5 Flash Lite](/models/google__gemini-3.5-flash-lite/)
- Google models [Google](/providers/google/)
- Where this model appears [wire](/wire/)
- Side by side [MiMo-V2.5-Pro](/compare/google__gemma-4-31b-it--vs--xiaomi__mimo-v2.5-pro/)
- Nearby on the capability axis [Claude Sonnet 5](/models/anthropic__claude-sonnet-5/)
- Same vendor [Gemma 4 26B A4B](/models/google__gemma-4-26b-a4b-it/)
- Where this model appears [self-host](/self-host/)
- Nearby on the capability axis [GPT-6 Astra](/models/openai__gpt-6-astra/)

## Badge

The verdict as one SVG, for a README or a docs page. It states this model's dominance status at the balanced workload on the LMArena lens, with the date it was computed, and it is rebuilt with the catalogue — so it changes when the verdict changes, including to one you would rather it did not.

 Markdown Copy
 [![Gemma 4 31B — Undominated.ai dominance verdict](https://undominated.ai/badge/google__gemma-4-31b-it.svg)](https://undominated.ai/models/google__gemma-4-31b-it/)
 HTML Copy
 <a href="https://undominated.ai/models/google__gemma-4-31b-it/"><img src="https://undominated.ai/badge/google__gemma-4-31b-it.svg" alt="Gemma 4 31B — Undominated.ai dominance verdict" height="36"></a>

Direct file: [https://undominated.ai/badge/google__gemma-4-31b-it.svg](https://undominated.ai/badge/google__gemma-4-31b-it.svg) — a static SVG, written by scripts/build-badges.mjs on every build, so the copy you embed is never older than the last deploy.

## Continue your investigation

 - [Explore model families](/families/)
- [Compare a shortlist](/compare/)
- [Check a model](/check/)
- [Understand the evidence](/methodology/)
