Reasoning tokens are billed separately at $3.75/M, on
top of output. Its share of your bill depends on the reasoning tokens used.
Who sells it
The same weights, different shops. Cheapest is not like-for-like when serving precision differs.
6 provider offers · rates, limits and conditions
Rates are USD per million tokens. Endpoint terms can differ even when the quoted price and precision match. A listed parameter is a provider declaration, not a task-success test.
Seller
Input
Output
Precision
Uptime (1d)
Seller
Google AI Studiogoogle-ai-studio/flex Endpoint terms
Cached input /M
$0.037
Context limit
1,048,576 tokens
Output limit
65,536 tokens
Tools
Listed by endpoint
Reasoning
Listed by endpoint
Promotional discount
50% · already included in these rates
Uptime · last 30 minutes
100.00%
Input$0.375
Output$1.88
Precisionundeclared
Uptime (1d)99.99%
Seller
Googlegoogle-vertex/global/flex Endpoint terms
Cached input /M
$0.037
Context limit
1,048,576 tokens
Output limit
65,536 tokens
Tools
Listed by endpoint
Reasoning
Listed by endpoint
Promotional discount
50% · already included in these rates
Uptime · last 30 minutes
Unknown
Input$0.375
Output$1.88
Precisionundeclared
Uptime (1d)97.76%
Seller
Google AI Studiogoogle-ai-studio Endpoint terms
Cached input /M
$0.075
Context limit
1,048,576 tokens
Output limit
65,536 tokens
Tools
Listed by endpoint
Reasoning
Listed by endpoint
Promotional discount
50% · already included in these rates
Uptime · last 30 minutes
99.93%
Input$0.750
Output$3.75
Precisionundeclared
Uptime (1d)99.93%
Seller
Googlegoogle-vertex/global Endpoint terms
Cached input /M
$0.075
Context limit
1,048,576 tokens
Output limit
65,536 tokens
Tools
Listed by endpoint
Reasoning
Listed by endpoint
Promotional discount
50% · already included in these rates
Uptime · last 30 minutes
92.67%
Input$0.750
Output$3.75
Precisionundeclared
Uptime (1d)94.79%
Seller
Google AI Studiogoogle-ai-studio/priority Endpoint terms
Gemini 3.8 Flash is Google's most intelligent Flash model with significant gains from 3.7 Flash across software engineering, agentic tasks, and multi-step reasoning.
Questions this page answers
What does Gemini 3.8 Flash cost?
$0.750/M in, $3.75/M out in this catalogue, as of the fetch date on this page. That is the API row, not a subscription.
Does Gemini 3.8 Flash have an independent quality score?
Gemini 3.8 Flash has an LMArena score in this catalogue.
What beats Gemini 3.8 Flash?
Nothing is both better and cheaper than Gemini 3.8 Flash.
Does this page use Artificial Analysis scores for Gemini 3.8 Flash?
No. Artificial Analysis figures are not published here. Quality on this page is LMArena Elo where a score exists; otherwise the row is unrated.
MODEL MONUMENT
Save or share this proof
Preview and more sharing options
Google: Gemini 3.8 Flash. Balanced workload. Frontier. Evidence as of 23 Sept 2026.
The verdict as one SVG, for a README or a docs page. It states this model's dominance
status at the balanced workload on the LMArena lens, with the date it was computed, and
it is rebuilt with the catalogue — so it changes when the verdict changes, including to
one you would rather it did not.