---
title: "Compare alternatives to Gemma 4 31B · Undominated.ai"
canonical: https://undominated.ai/alternatives/google__gemma-4-31b-it/
description: "3 alternatives score at least as high as Gemma 4 31B and cost no more under at least one listed workload. 2 preserve recorded capabilities; 1 have named trade-offs."
---

# Compare alternatives to Gemma 4 31B · Undominated.ai

> 3 alternatives score at least as high as Gemma 4 31B and cost no more under at least one listed workload. 2 preserve recorded capabilities; 1 have named trade-offs.

# Alternatives to Gemma 4 31B

Google · LMArena 1443.4 · $0.205/M on a balanced workload · prices as of 2026-10-06

 [Compare models](https://undominated.ai/compare/) Comparison context

Choose up to three standard models in Compare.

Requirements and prompt length apply in Compare. This page keeps its stated price basis and workload controls.

## 2 models are higher-scoring or cheaper, with neither dimension worse.

Of the 144 models carrying both an independent score and a published price at the same delivery mode, **3** score at least as high as Gemma 4 31B and cost no more under at least one listed workload, with at least one of those dimensions improved. **2** match or beat its recorded context, maximum output, input modes, tool use and extended reasoning. **1** has a recorded capability loss, named on every row.

Dominance is not a property of a model. It is a property of a model, a capability lens and a workload, and all three are named on every row. Scores are LMArena Elo, used under CC BY 4.0 from the official dataset; prices are blended per million tokens. A higher score is not a drop-in replacement.

This page is built only while a candidate improves score or price without worsening the other under a listed workload. When that comparison no longer holds, it is not generated.

## Where it holds

| Workload | Gemma 4 31B $/M | Recorded capabilities preserved | Recorded capability losses |
| --- | --- | --- | --- |
| Balanced | $0.205 | 1 | 1 |
| Summarise | $0.153 | 2 | 1 |
| Chat | $0.244 | 1 | 0 |
| Code gen | $0.296 | 1 | 0 |
| Agentic | $0.179 | 2 | 1 |

Workload mixes are defined on the [methodology page](/methodology/). Cached input is priced at the cached rate, and reasoning tokens at the worse of the reasoning and output rates.

## What you would be replacing

| Intelligence | 1443.4 |
| --- | --- |
| Context window | 262K |
| Max output | 16K |
| Input modes | image, text, video |
| Tool use | yes |
| Extended reasoning | yes |

Preserving recorded capabilities requires matching or beating these fields. [Full record for Gemma 4 31B](/models/google__gemma-4-31b-it/), or [check Gemma 4 31B against the whole catalogue](/check/google__gemma-4-31b-it/).

## Recorded capabilities preserved 2

Each option has a higher score or a lower price, with neither dimension worse under the listed workloads. It matches or beats Gemma 4 31B on recorded context, maximum output, input modes, tool use and extended reasoning.

| # | Model | Intelligence | $/M | Holds under |
| --- | --- | --- | --- | --- |
| 1 | [MiMo-V2.6-Flash](/models/xiaomi__mimo-v2.6-flash/) Xiaomi | 1456.4 +13 | $0.175 | Balanced −15% Summarise −29% Chat −37% Code gen −28% Agentic −56% |
| 2 | [GLM 5.3 Flash](/models/z-ai__glm-5.3-flash/) Z.ai | 1469.6 +26.2 | $0.133 summarise | Summarise −13% Agentic −27% |

## Alternatives with recorded capability losses 1

These improve score or price without worsening the other under the listed workloads, but reduce at least one recorded capability. Read the last column before switching.

| Model | Intelligence | $/M | Holds under | Recorded capability loss |
| --- | --- | --- | --- | --- |
| [DeepSeek V4.1 Flash](/models/deepseek__deepseek-v4.1-flash/) DeepSeek | 1462.5 +19.1 | $0.182 | Balanced −11% Summarise −56% Agentic −41% | no video input |

## What this compares, and what it leaves out

 - Quality is LMArena Elo, used under CC BY 4.0 from the official dataset. The 95% confidence interval on a difference between two scores is about ±10.47 points, so a gap smaller than that is marked *tie* rather than an improvement — see [significance bands](/significance/).
 - 204 further models at this delivery mode carry a price but no independent score. They are absent from the comparison in both directions — unrated is not a zero, and an unmeasured model is neither an alternative nor a worse buy.
 - Retired models are never offered as an alternative, and a model only competes against its own delivery mode: batch trades latency for price, so it is not a like-for-like swap.
 - Nothing here measures latency, throughput, rate limits or how a model behaves on your prompts. Two models with the same index score are not interchangeable.
 - Models with no higher-scoring or cheaper alternative that is no worse on the other axis are on the [value frontier](/frontier/). Models with such an alternative are [listed here](/alternatives/).

## Continue your investigation

 - [Check requirements](/check/)
- [Compare exact differences](/compare/)
- [Set a quality floor](/cheapest-at/)
- [Inspect tier boundaries](/cliffs/)
