MODEL PROOF

Gemma 4 31B

FRONTIER

Nothing is both better and cheaper under this workload.

$0.152 / 1M tokens Balanced · 3 tokens in per 1 out

Serving precision differs between offers.

Independent LMArena score
1,441.7 ± 7.54 · 5,894 votes
Context window
262.14K
Maximum output
16.38K
Input modalities
image, text, video
Output modalities
text
Published input price
$0.090 / 1M tokens
Published output price
$0.340 / 1M tokens
Pricing kind
fixed

Inspect complete billing conditions and endpoint terms below.

Every price dimension · USD per million tokens
Input$0.090/M
Output$0.340/M
Cached input$0.050/M
Cache writeUnknown/M
Cache write 1hUnknown/M
ReasoningUnknown/M

Who sells it

The same weights, different shops. Cheapest is not like-for-like when serving precision differs.

Headline cheapest is a lower precision. Like-for-like at the best declared precision is $0.140/M from Crusoe.

13 provider offers · rates, limits and conditions

Rates are USD per million tokens. Endpoint terms can differ even when the quoted price and precision match. A listed parameter is a provider declaration, not a task-success test.

SellerInputOutputPrecisionUptime (1d)
DeepInfradeepinfra/turbo
Endpoint terms
Cached input /M
$0.050
Context limit
262,144 tokens
Output limit
16,384 tokens
Tools
Not listed by endpoint
Reasoning
Listed by endpoint
Promotional discount
None reported
Uptime · last 30 minutes
98.83%
$0.090 $0.340 fp4 98.42%
CoreWeavecoreweave/fp4
Endpoint terms
Cached input /M
$0.100
Context limit
262,144 tokens
Output limit
235,929 tokens
Tools
Listed by endpoint
Reasoning
Listed by endpoint
Promotional discount
None reported
Uptime · last 30 minutes
96.77%
$0.100 $0.340 fp4 97.53%
Venicevenice/fp4
Endpoint terms
Cached input /M
$0.090
Context limit
256,000 tokens
Output limit
8,192 tokens
Tools
Listed by endpoint
Reasoning
Listed by endpoint
Promotional discount
None reported
Uptime · last 30 minutes
96.15%
$0.120 $0.360 fp4 97.66%
Chuteschutes/fp4
Endpoint terms
Cached input /M
$0.012
Context limit
131,072 tokens
Output limit
65,536 tokens
Tools
Listed by endpoint
Reasoning
Listed by endpoint
Promotional discount
None reported
Uptime · last 30 minutes
95.09%
$0.120 $0.370 fp4 94.23%
DeepInfradeepinfra/fp8
Endpoint terms
Cached input /M
Unknown
Context limit
262,144 tokens
Output limit
16,384 tokens
Tools
Listed by endpoint
Reasoning
Listed by endpoint
Promotional discount
None reported
Uptime · last 30 minutes
96.55%
$0.130 $0.380 fp8 96.82%
Crusoecrusoe/bf16
Endpoint terms
Cached input /M
$0.140
Context limit
262,144 tokens
Output limit
262,141 tokens
Tools
Listed by endpoint
Reasoning
Listed by endpoint
Promotional discount
None reported
Uptime · last 30 minutes
99.76%
$0.140 $0.400 bf16 98.25%
Novitanovita/bf16
Endpoint terms
Cached input /M
Unknown
Context limit
262,144 tokens
Output limit
131,072 tokens
Tools
Listed by endpoint
Reasoning
Listed by endpoint
Promotional discount
None reported
Uptime · last 30 minutes
90.45%
$0.140 $0.400 bf16 93.77%
Friendlifriendli
Endpoint terms
Cached input /M
Unknown
Context limit
262,144 tokens
Output limit
8,192 tokens
Tools
Listed by endpoint
Reasoning
Listed by endpoint
Promotional discount
None reported
Uptime · last 30 minutes
99.91%
$0.140 $0.400 undeclared 99.63%
Parasailparasail/fp8
Endpoint terms
Cached input /M
$0.060
Context limit
262,144 tokens
Output limit
235,929 tokens
Tools
Listed by endpoint
Reasoning
Listed by endpoint
Promotional discount
None reported
Uptime · last 30 minutes
98.48%
$0.150 $0.400 fp8 99.45%
DeepInfradeepinfra/ultra
Endpoint terms
Cached input /M
Unknown
Context limit
131,072 tokens
Output limit
8,192 tokens
Tools
Listed by endpoint
Reasoning
Listed by endpoint
Promotional discount
None reported
Uptime · last 30 minutes
2.73%
$0.270 $0.760 fp8 27.18%
SambaNovasambanova
Endpoint terms
Cached input /M
Unknown
Context limit
262,144 tokens
Output limit
235,929 tokens
Tools
Listed by endpoint
Reasoning
Listed by endpoint
Promotional discount
None reported
Uptime · last 30 minutes
86.62%
$0.380 $1.15 undeclared 96.92%
SiliconFlowsiliconflow/fp8
Endpoint terms
Cached input /M
$0.250
Context limit
262,144 tokens
Output limit
235,929 tokens
Tools
Listed by endpoint
Reasoning
Listed by endpoint
Promotional discount
None reported
Uptime · last 30 minutes
62.60%
$0.750 $1.00 fp8 80.08%
ModelRunmodelrun/fp4
Endpoint terms
Cached input /M
$0.750
Context limit
262,144 tokens
Output limit
235,929 tokens
Tools
Listed by endpoint
Reasoning
Listed by endpoint
Promotional discount
None reported
Uptime · last 30 minutes
100.00%
$0.750 $1.00 fp4 99.99%

Independent scores

LMArena Elo1441.7default

LMArena Elo under CC BY 4.0. Absence of another board is not a score of zero.

Capability

Context window262K
Max output16K
Input modesimage, text, video
Tool useyes
Reasoningoptional
Open weightsyes

Provenance

Price sourceopenrouter.ai
Fetched2026-09-23
Quality dataverified
Cross-checkedvendor page

Gemma 4 31B Instruct is Google DeepMind's 30.7B dense multimodal model supporting text and image input with text output. Features a 256K token context window, configurable thinking/reasoning mode, native function...

Questions this page answers

What does Gemma 4 31B cost?

$0.090/M in, $0.340/M out in this catalogue, as of the fetch date on this page. That is the API row, not a subscription.

Does Gemma 4 31B have an independent quality score?

Gemma 4 31B has an LMArena score in this catalogue.

What beats Gemma 4 31B?

Nothing is both better and cheaper than Gemma 4 31B.

Does this page use Artificial Analysis scores for Gemma 4 31B?

No. Artificial Analysis figures are not published here. Quality on this page is LMArena Elo where a score exists; otherwise the row is unrated.

Is the cheapest Gemma 4 31B endpoint the same product?

Headline cheapest is a lower precision. Like-for-like at the best declared precision is $0.140/M from Crusoe.

MODEL MONUMENT

Save or share this proof

Badge

The verdict as one SVG, for a README or a docs page. It states this model's dominance status at the balanced workload on the LMArena lens, with the date it was computed, and it is rebuilt with the catalogue — so it changes when the verdict changes, including to one you would rather it did not.

Gemma 4 31B — Undominated.ai dominance verdict

Markdown
[![Gemma 4 31B — Undominated.ai dominance verdict](https://undominated.ai/badge/google__gemma-4-31b-it.svg)](https://undominated.ai/models/google__gemma-4-31b-it/)
HTML
<a href="https://undominated.ai/models/google__gemma-4-31b-it/"><img src="https://undominated.ai/badge/google__gemma-4-31b-it.svg" alt="Gemma 4 31B — Undominated.ai dominance verdict" height="36"></a>

Direct file: https://undominated.ai/badge/google__gemma-4-31b-it.svg — a static SVG, written by scripts/build-badges.mjs on every build, so the copy you embed is never older than the last deploy.

Evidence & Ask