---
title: "Google: Gemma 4 31B — price, capability and what beats it · Undominated.ai"
canonical: https://undominated.ai/models/google__gemma-4-31b-it/
description: "Google: Gemma 4 31B: $0.100/M in, $0.340/M out. Independent capability scores, context-tier pricing, and the models that are both better and cheaper."
---

# Google: Gemma 4 31B — price, capability and what beats it · Undominated.ai

> Google: Gemma 4 31B: $0.100/M in, $0.340/M out. Independent capability scores, context-tier pricing, and the models that are both better and cheaper.

[Leaderboard](/) / Google

# Gemma 4 31B

Google · released 2026-04-02 · gemma

 $0.160 per million tokens, balanced Balanced Summarise Chat Code gen Agentic

## DeepSeek V4 Flash 0731 scores higher and costs less — but you would give something up.

**+22.1** on the capability index and **34% cheaper** ($0.105/M against $0.160/M). What you lose:

 - no image, video input

Capability scores do not measure context length, output ceiling or which inputs a model accepts, so a higher score does not mean a drop-in replacement.

The same model via **free** is 100% cheaper (free/M) — same weights, different latency.

### Every price dimension

| Input | $0.100 /M |
| --- | --- |
| Output | $0.340 /M |
| Cached input | $0.100 /M |

### Independent scores

| Intelligence | 29.7 |
| --- | --- |
| Coding | 43.4 |
| Agentic | 14.4 |
| LMArena Elo | 1441.7 default |

Sources are listed separately rather than averaged. Across the 62 models both have scored they correlate at r = 0.806 — close agreement overall, but four models rank very differently between them, and a blended score would hide exactly those.

### Capability

| Context window | 262K |
| --- | --- |
| Max output | 262K |
| Input modes | image, text, video |
| Tool use | yes |
| Reasoning | optional |
| Open weights | yes |

### Provenance

| Price source | openrouter.ai |
| --- | --- |
| Fetched | 2026-08-24 |
| Quality data | verified |
| Cross-checked | vendor page |

Gemma 4 31B Instruct is Google DeepMind's 30.7B dense multimodal model supporting text and image input with text output. Features a 256K token context window, configurable thinking/reasoning mode, native function...
