Price from Io Net: the cheapest offer that is serving, at standard delivery and at a declared precision that is not 4-bit (bf16), at its standard rate. The maker’s own precision is not known.
Cheaper right now: $0.059 in · $0.255 out per million tokens at Io Net, a promotion of 15%. The price on this page is Io Net’s standard rate: the ranking uses the standard rate, never a promotion.
Every price dimension · USD per million tokens
Input
$0.069/M
Output
$0.300/M
Cached input
$0.029/M
Cache write
Unknown/M
Cache write 1h
Unknown/M
Reasoning
Unknown/M
Who sells it
The same weights, different shops. Cheapest is not like-for-like when serving precision differs.
Headline cheapest is a lower precision. Like-for-like at the best declared precision is $0.060/M from DekaLLM.
13 provider offers · rates, limits and conditions
Rates are USD per million tokens. Endpoint terms can differ even when the quoted price and precision match. A listed parameter is a provider declaration, not a task-success test.
Seller
Input
Output
Precision
Uptime (1d)
Seller
Darkbloomdarkbloom Endpoint terms
Cached input /M
$0.021
Context limit
131,072 tokens
Output limit
32,768 tokens
Tools
Listed by endpoint
Reasoning
Listed by endpoint
Promotional discount
None reported
Uptime · last 30 minutes
100.00%
Input$0.042
Output$0.220
Precisionundeclared
Uptime (1d)99.99%
Seller
Io Netio-net/bf16 Endpoint terms
Cached input /M
$0.025
Context limit
262,142 tokens
Output limit
235,927 tokens
Tools
Listed by endpoint
Reasoning
Listed by endpoint
Promotional discount
15% · already included in these rates
Uptime · last 30 minutes
99.94%
Input$0.059
Output$0.255
Precisionbf16
Uptime (1d)99.75%
Seller
DekaLLMdekallm/bf16 Endpoint terms
Cached input /M
$0.040
Context limit
262,144 tokens
Output limit
235,929 tokens
Tools
Listed by endpoint
Reasoning
Listed by endpoint
Promotional discount
None reported
Uptime · last 30 minutes
99.82%
Input$0.060
Output$0.330
Precisionbf16
Uptime (1d)99.81%
Seller
DeepInfradeepinfra/fp8 Endpoint terms
Cached input /M
Unknown
Context limit
262,144 tokens
Output limit
16,384 tokens
Tools
Listed by endpoint
Reasoning
Listed by endpoint
Promotional discount
None reported
Uptime · last 30 minutes
99.80%
Input$0.070
Output$0.340
Precisionfp8
Uptime (1d)99.77%
Seller
Makoramakora Endpoint terms
Cached input /M
$0.032
Context limit
256,000 tokens
Output limit
128,000 tokens
Tools
Listed by endpoint
Reasoning
Listed by endpoint
Promotional discount
None reported
Uptime · last 30 minutes
99.83%
Input$0.080
Output$0.320
Precisionundeclared
Uptime (1d)98.52%
Seller
NextBitnextbit/bf16 Endpoint terms
Cached input /M
$0.050
Context limit
262,144 tokens
Output limit
235,929 tokens
Tools
Listed by endpoint
Reasoning
Listed by endpoint
Promotional discount
None reported
Uptime · last 30 minutes
99.75%
Input$0.090
Output$0.300
Precisionbf16
Uptime (1d)99.93%
Seller
CoreWeavecoreweave/bf16 Endpoint terms
Cached input /M
$0.050
Context limit
262,144 tokens
Output limit
235,929 tokens
Tools
Not listed by endpoint
Reasoning
Listed by endpoint
Promotional discount
None reported
Uptime · last 30 minutes
99.99%
Input$0.100
Output$0.300
Precisionbf16
Uptime (1d)99.97%
Seller
Cloudflarecloudflare Endpoint terms
Cached input /M
$0.050
Context limit
256,000 tokens
Output limit
230,400 tokens
Tools
Listed by endpoint
Reasoning
Listed by endpoint
Promotional discount
None reported
Uptime · last 30 minutes
99.89%
Input$0.100
Output$0.300
Precisionundeclared
Uptime (1d)99.86%
Seller
Venicevenice/bf16 Endpoint terms
Cached input /M
$0.050
Context limit
256,000 tokens
Output limit
8,192 tokens
Tools
Listed by endpoint
Reasoning
Listed by endpoint
Promotional discount
None reported
Uptime · last 30 minutes
98.18%
Input$0.130
Output$0.400
Precisionbf16
Uptime (1d)98.84%
Seller
Parasailparasail/bf16 Endpoint terms
Cached input /M
$0.050
Context limit
262,144 tokens
Output limit
235,929 tokens
Tools
Not listed by endpoint
Reasoning
Listed by endpoint
Promotional discount
None reported
Uptime · last 30 minutes
98.27%
Input$0.130
Output$0.400
Precisionbf16
Uptime (1d)99.15%
Seller
Novitanovita/bf16 Endpoint terms
Cached input /M
Unknown
Context limit
262,144 tokens
Output limit
131,072 tokens
Tools
Listed by endpoint
Reasoning
Listed by endpoint
Promotional discount
None reported
Uptime · last 30 minutes
99.85%
Input$0.130
Output$0.400
Precisionbf16
Uptime (1d)99.81%
Seller
SiliconFlowsiliconflow/fp8 Endpoint terms
Cached input /M
$0.050
Context limit
262,144 tokens
Output limit
235,929 tokens
Tools
Not listed by endpoint
Reasoning
Listed by endpoint
Promotional discount
None reported
Uptime · last 30 minutes
96.52%
Input$0.140
Output$0.400
Precisionfp8
Uptime (1d)95.41%
Seller
Googlegoogle-vertex/global Endpoint terms
Cached input /M
Unknown
Context limit
262,144 tokens
Output limit
235,929 tokens
Tools
Listed by endpoint
Reasoning
Listed by endpoint
Promotional discount
None reported
Uptime · last 30 minutes
98.47%
Input$0.150
Output$0.600
Precisionundeclared
Uptime (1d)97.49%
Independent scores
LMArena Elo
1433.9default
LMArena Elo under CC BY 4.0. Absence of another board is not a score of zero.
Gemma 4 26B A4B IT is an instruction-tuned Mixture-of-Experts (MoE) model from Google DeepMind. Despite 25.2B total parameters, only 3.8B activate per token during inference — delivering near-31B quality at...
Questions this page answers
What does Gemma 4 26B A4B cost?
$0.069/M in, $0.300/M out in this catalogue, as of the fetch date on this page. That is the API row, not a subscription.
Does Gemma 4 26B A4B have an independent quality score?
Gemma 4 26B A4B has an LMArena score in this catalogue.
What beats Gemma 4 26B A4B ?
Nothing is both better and cheaper than Gemma 4 26B A4B .
Does this page use Artificial Analysis scores for Gemma 4 26B A4B ?
No. Artificial Analysis figures are not published here. Quality on this page is LMArena Elo where a score exists; otherwise the row is unrated.
Is the cheapest Gemma 4 26B A4B endpoint the same product?
Headline cheapest is a lower precision. Like-for-like at the best declared precision is $0.060/M from DekaLLM.
MODEL MONUMENT
Save or share this proof
Preview and more sharing options
Google: Gemma 4 26B A4B . Balanced workload. Frontier. Evidence as of 06 Oct 2026.
The verdict as one SVG, for a README or a docs page. It states this model's dominance
status at the balanced workload on the LMArena lens, with the date it was computed, and
it is rebuilt with the catalogue — so it changes when the verdict changes, including to
one you would rather it did not.