The same weights, different shops. Cheapest is not like-for-like when serving precision differs.
Headline cheapest is a lower precision. Like-for-like at the best declared precision is $0.140/M from Crusoe.
13 provider offers · rates, limits and conditions
Rates are USD per million tokens. Endpoint terms can differ even when the quoted price and precision match. A listed parameter is a provider declaration, not a task-success test.
Seller
Input
Output
Precision
Uptime (1d)
Seller
DeepInfradeepinfra/turbo Endpoint terms
Cached input /M
$0.050
Context limit
262,144 tokens
Output limit
16,384 tokens
Tools
Not listed by endpoint
Reasoning
Listed by endpoint
Promotional discount
None reported
Uptime · last 30 minutes
98.83%
Input$0.090
Output$0.340
Precisionfp4
Uptime (1d)98.42%
Seller
CoreWeavecoreweave/fp4 Endpoint terms
Cached input /M
$0.100
Context limit
262,144 tokens
Output limit
235,929 tokens
Tools
Listed by endpoint
Reasoning
Listed by endpoint
Promotional discount
None reported
Uptime · last 30 minutes
96.77%
Input$0.100
Output$0.340
Precisionfp4
Uptime (1d)97.53%
Seller
Venicevenice/fp4 Endpoint terms
Cached input /M
$0.090
Context limit
256,000 tokens
Output limit
8,192 tokens
Tools
Listed by endpoint
Reasoning
Listed by endpoint
Promotional discount
None reported
Uptime · last 30 minutes
96.15%
Input$0.120
Output$0.360
Precisionfp4
Uptime (1d)97.66%
Seller
Chuteschutes/fp4 Endpoint terms
Cached input /M
$0.012
Context limit
131,072 tokens
Output limit
65,536 tokens
Tools
Listed by endpoint
Reasoning
Listed by endpoint
Promotional discount
None reported
Uptime · last 30 minutes
95.09%
Input$0.120
Output$0.370
Precisionfp4
Uptime (1d)94.23%
Seller
DeepInfradeepinfra/fp8 Endpoint terms
Cached input /M
Unknown
Context limit
262,144 tokens
Output limit
16,384 tokens
Tools
Listed by endpoint
Reasoning
Listed by endpoint
Promotional discount
None reported
Uptime · last 30 minutes
96.55%
Input$0.130
Output$0.380
Precisionfp8
Uptime (1d)96.82%
Seller
Crusoecrusoe/bf16 Endpoint terms
Cached input /M
$0.140
Context limit
262,144 tokens
Output limit
262,141 tokens
Tools
Listed by endpoint
Reasoning
Listed by endpoint
Promotional discount
None reported
Uptime · last 30 minutes
99.76%
Input$0.140
Output$0.400
Precisionbf16
Uptime (1d)98.25%
Seller
Novitanovita/bf16 Endpoint terms
Cached input /M
Unknown
Context limit
262,144 tokens
Output limit
131,072 tokens
Tools
Listed by endpoint
Reasoning
Listed by endpoint
Promotional discount
None reported
Uptime · last 30 minutes
90.45%
Input$0.140
Output$0.400
Precisionbf16
Uptime (1d)93.77%
Seller
Friendlifriendli Endpoint terms
Cached input /M
Unknown
Context limit
262,144 tokens
Output limit
8,192 tokens
Tools
Listed by endpoint
Reasoning
Listed by endpoint
Promotional discount
None reported
Uptime · last 30 minutes
99.91%
Input$0.140
Output$0.400
Precisionundeclared
Uptime (1d)99.63%
Seller
Parasailparasail/fp8 Endpoint terms
Cached input /M
$0.060
Context limit
262,144 tokens
Output limit
235,929 tokens
Tools
Listed by endpoint
Reasoning
Listed by endpoint
Promotional discount
None reported
Uptime · last 30 minutes
98.48%
Input$0.150
Output$0.400
Precisionfp8
Uptime (1d)99.45%
Seller
DeepInfradeepinfra/ultra Endpoint terms
Cached input /M
Unknown
Context limit
131,072 tokens
Output limit
8,192 tokens
Tools
Listed by endpoint
Reasoning
Listed by endpoint
Promotional discount
None reported
Uptime · last 30 minutes
2.73%
Input$0.270
Output$0.760
Precisionfp8
Uptime (1d)27.18%
Seller
SambaNovasambanova Endpoint terms
Cached input /M
Unknown
Context limit
262,144 tokens
Output limit
235,929 tokens
Tools
Listed by endpoint
Reasoning
Listed by endpoint
Promotional discount
None reported
Uptime · last 30 minutes
86.62%
Input$0.380
Output$1.15
Precisionundeclared
Uptime (1d)96.92%
Seller
SiliconFlowsiliconflow/fp8 Endpoint terms
Cached input /M
$0.250
Context limit
262,144 tokens
Output limit
235,929 tokens
Tools
Listed by endpoint
Reasoning
Listed by endpoint
Promotional discount
None reported
Uptime · last 30 minutes
62.60%
Input$0.750
Output$1.00
Precisionfp8
Uptime (1d)80.08%
Seller
ModelRunmodelrun/fp4 Endpoint terms
Cached input /M
$0.750
Context limit
262,144 tokens
Output limit
235,929 tokens
Tools
Listed by endpoint
Reasoning
Listed by endpoint
Promotional discount
None reported
Uptime · last 30 minutes
100.00%
Input$0.750
Output$1.00
Precisionfp4
Uptime (1d)99.99%
Independent scores
LMArena Elo
1441.7default
LMArena Elo under CC BY 4.0. Absence of another board is not a score of zero.
Gemma 4 31B Instruct is Google DeepMind's 30.7B dense multimodal model supporting text and image input with text output. Features a 256K token context window, configurable thinking/reasoning mode, native function...
Questions this page answers
What does Gemma 4 31B cost?
$0.090/M in, $0.340/M out in this catalogue, as of the fetch date on this page. That is the API row, not a subscription.
Does Gemma 4 31B have an independent quality score?
Gemma 4 31B has an LMArena score in this catalogue.
What beats Gemma 4 31B?
Nothing is both better and cheaper than Gemma 4 31B.
Does this page use Artificial Analysis scores for Gemma 4 31B?
No. Artificial Analysis figures are not published here. Quality on this page is LMArena Elo where a score exists; otherwise the row is unrated.
Is the cheapest Gemma 4 31B endpoint the same product?
Headline cheapest is a lower precision. Like-for-like at the best declared precision is $0.140/M from Crusoe.
MODEL MONUMENT
Save or share this proof
Preview and more sharing options
Google: Gemma 4 31B. Balanced workload. Frontier. Evidence as of 23 Sept 2026.
The verdict as one SVG, for a README or a docs page. It states this model's dominance
status at the balanced workload on the LMArena lens, with the date it was computed, and
it is rebuilt with the catalogue — so it changes when the verdict changes, including to
one you would rather it did not.