Price from Crusoe: the cheapest offer that is serving, at standard delivery and at a declared precision that is not 4-bit (bf16), at its standard rate. The maker’s own precision is not known.
Cheaper right now: $0.153/M at DeepInfra — delivery tier (turbo) · 4-bit It does not pass the like-for-like test, so it never sets a rank.
Every price dimension · USD per million tokens
Input
$0.140/M
Output
$0.400/M
Cached input
$0.140/M
Cache write
Unknown/M
Cache write 1h
Unknown/M
Reasoning
Unknown/M
Who sells it
The same weights, different shops. Cheapest is not like-for-like when serving precision differs.
Headline cheapest is a lower precision. Like-for-like at the best declared precision is $0.140/M from Crusoe.
13 provider offers · rates, limits and conditions
Rates are USD per million tokens. Endpoint terms can differ even when the quoted price and precision match. A listed parameter is a provider declaration, not a task-success test.
Seller
Input
Output
Precision
Uptime (1d)
Seller
DeepInfradeepinfra/turbo Endpoint terms
Cached input /M
$0.050
Context limit
262,144 tokens
Output limit
16,384 tokens
Tools
Not listed by endpoint
Reasoning
Listed by endpoint
Promotional discount
None reported
Uptime · last 30 minutes
98.90%
Input$0.090
Output$0.340
Precisionfp4
Uptime (1d)99.29%
Seller
CoreWeavecoreweave/fp4 Endpoint terms
Cached input /M
$0.100
Context limit
262,144 tokens
Output limit
235,929 tokens
Tools
Listed by endpoint
Reasoning
Listed by endpoint
Promotional discount
None reported
Uptime · last 30 minutes
99.55%
Input$0.100
Output$0.340
Precisionfp4
Uptime (1d)99.24%
Seller
Venicevenice/fp4 Endpoint terms
Cached input /M
$0.090
Context limit
256,000 tokens
Output limit
8,192 tokens
Tools
Listed by endpoint
Reasoning
Listed by endpoint
Promotional discount
None reported
Uptime · last 30 minutes
98.14%
Input$0.120
Output$0.360
Precisionfp4
Uptime (1d)99.41%
Seller
Chuteschutes/fp4 Endpoint terms
Cached input /M
$0.012
Context limit
131,072 tokens
Output limit
65,536 tokens
Tools
Listed by endpoint
Reasoning
Listed by endpoint
Promotional discount
None reported
Uptime · last 30 minutes
95.93%
Input$0.120
Output$0.370
Precisionfp4
Uptime (1d)94.20%
Seller
Crusoecrusoe/bf16 Endpoint terms
Cached input /M
$0.140
Context limit
262,144 tokens
Output limit
262,141 tokens
Tools
Listed by endpoint
Reasoning
Listed by endpoint
Promotional discount
None reported
Uptime · last 30 minutes
99.89%
Input$0.140
Output$0.400
Precisionbf16
Uptime (1d)99.48%
Seller
Novitanovita/bf16 Endpoint terms
Cached input /M
Unknown
Context limit
262,144 tokens
Output limit
131,072 tokens
Tools
Listed by endpoint
Reasoning
Listed by endpoint
Promotional discount
None reported
Uptime · last 30 minutes
81.87%
Input$0.140
Output$0.400
Precisionbf16
Uptime (1d)82.80%
Seller
Friendlifriendli Endpoint terms
Cached input /M
Unknown
Context limit
262,144 tokens
Output limit
8,192 tokens
Tools
Listed by endpoint
Reasoning
Listed by endpoint
Promotional discount
None reported
Uptime · last 30 minutes
99.93%
Input$0.140
Output$0.400
Precisionundeclared
Uptime (1d)99.59%
Seller
Parasailparasail/fp8 Endpoint terms
Cached input /M
$0.060
Context limit
262,144 tokens
Output limit
235,929 tokens
Tools
Listed by endpoint
Reasoning
Listed by endpoint
Promotional discount
None reported
Uptime · last 30 minutes
98.96%
Input$0.150
Output$0.400
Precisionfp8
Uptime (1d)99.33%
Seller
DeepInfradeepinfra/fp8 Endpoint terms
Cached input /M
Unknown
Context limit
262,144 tokens
Output limit
16,384 tokens
Tools
Listed by endpoint
Reasoning
Listed by endpoint
Promotional discount
None reported
Uptime · last 30 minutes
99.52%
Input$0.200
Output$0.400
Precisionfp8
Uptime (1d)98.77%
Seller
Io Netio-net Endpoint terms
Cached input /M
$0.181
Context limit
262,144 tokens
Output limit
16,384 tokens
Tools
Listed by endpoint
Reasoning
Listed by endpoint
Promotional discount
5% · already included in these rates
Uptime · last 30 minutes
99.35%
Input$0.361
Output$1.09
Precisionundeclared
Uptime (1d)99.28%
Seller
SambaNovasambanova Endpoint terms
Cached input /M
Unknown
Context limit
262,144 tokens
Output limit
235,929 tokens
Tools
Listed by endpoint
Reasoning
Listed by endpoint
Promotional discount
None reported
Uptime · last 30 minutes
97.95%
Input$0.380
Output$1.15
Precisionundeclared
Uptime (1d)98.55%
Seller
SiliconFlowsiliconflow/fp8 Endpoint terms
Cached input /M
$0.250
Context limit
262,144 tokens
Output limit
235,929 tokens
Tools
Listed by endpoint
Reasoning
Listed by endpoint
Promotional discount
None reported
Uptime · last 30 minutes
99.18%
Input$0.750
Output$1.00
Precisionfp8
Uptime (1d)76.52%
Seller
ModelRunmodelrun/fp4 Endpoint terms
Cached input /M
$0.200
Context limit
262,144 tokens
Output limit
235,929 tokens
Tools
Listed by endpoint
Reasoning
Listed by endpoint
Promotional discount
None reported
Uptime · last 30 minutes
99.36%
Input$0.750
Output$1.00
Precisionfp4
Uptime (1d)98.99%
Independent scores
LMArena Elo
1443.4default
LMArena Elo under CC BY 4.0. Absence of another board is not a score of zero.
Gemma 4 31B Instruct is Google DeepMind's 30.7B dense multimodal model supporting text and image input with text output. Features a 256K token context window, configurable thinking/reasoning mode, native function...
Questions this page answers
What does Gemma 4 31B cost?
$0.140/M in, $0.400/M out in this catalogue, as of the fetch date on this page. That is the API row, not a subscription.
Does Gemma 4 31B have an independent quality score?
Gemma 4 31B has an LMArena score in this catalogue.
What beats Gemma 4 31B?
MiMo-V2.6-Flash has an equal-or-higher measured score and an equal-or-lower price.
Does this page use Artificial Analysis scores for Gemma 4 31B?
No. Artificial Analysis figures are not published here. Quality on this page is LMArena Elo where a score exists; otherwise the row is unrated.
Is the cheapest Gemma 4 31B endpoint the same product?
Headline cheapest is a lower precision. Like-for-like at the best declared precision is $0.140/M from Crusoe.
MODEL MONUMENT
Save or share this proof
Preview and more sharing options
Google: Gemma 4 31B. Balanced workload. Better value option available. Evidence as of 06 Oct 2026.
The verdict as one SVG, for a README or a docs page. It states this model's dominance
status at the balanced workload on the LMArena lens, with the date it was computed, and
it is rebuilt with the catalogue — so it changes when the verdict changes, including to
one you would rather it did not.