Price from DeepInfra: the cheapest offer that is serving, at standard delivery and at a declared precision that is not 4-bit (fp8), at its standard rate. The maker’s own precision is not known.
Cheaper right now: $0.153/M at GMICloud — promotion 75% It does not pass the like-for-like test, so it never sets a rank.
Every price dimension · USD per million tokens
Input
$0.090/M
Output
$0.550/M
Cached input
Unknown/M
Cache write
Unknown/M
Cache write 1h
Unknown/M
Reasoning
Unknown/M
Who sells it
The same weights, different shops. Cheapest is not like-for-like when serving precision differs.
Headline cheapest is a lower precision. Like-for-like at the best declared precision is $0.090/M from DeepInfra.
10 provider offers · rates, limits and conditions
Rates are USD per million tokens. Endpoint terms can differ even when the quoted price and precision match. A listed parameter is a provider declaration, not a task-success test.
Seller
Input
Output
Precision
Uptime (1d)
Seller
GMICloudgmicloud/fp8 Endpoint terms
Cached input /M
$0.018
Context limit
262,144 tokens
Output limit
235,929 tokens
Tools
Listed by endpoint
Reasoning
Not listed by endpoint
Promotional discount
75% · already included in these rates
Uptime · last 30 minutes
99.67%
Input$0.088
Output$0.350
Precisionfp8
Uptime (1d)99.40%
Seller
DeepInfradeepinfra/fp8 Endpoint terms
Cached input /M
Unknown
Context limit
262,144 tokens
Output limit
16,384 tokens
Tools
Listed by endpoint
Reasoning
Not listed by endpoint
Promotional discount
None reported
Uptime · last 30 minutes
93.62%
Input$0.090
Output$0.550
Precisionfp8
Uptime (1d)95.46%
Seller
Novitanovita/fp8 Endpoint terms
Cached input /M
Unknown
Context limit
131,072 tokens
Output limit
16,384 tokens
Tools
Listed by endpoint
Reasoning
Not listed by endpoint
Promotional discount
None reported
Uptime · last 30 minutes
94.61%
Input$0.090
Output$0.580
Precisionfp8
Uptime (1d)96.01%
Seller
Parasailparasail/fp8 Endpoint terms
Cached input /M
$0.050
Context limit
131,072 tokens
Output limit
117,964 tokens
Tools
Listed by endpoint
Reasoning
Not listed by endpoint
Promotional discount
None reported
Uptime · last 30 minutes
99.93%
Input$0.140
Output$0.800
Precisionfp8
Uptime (1d)99.90%
Seller
Alibabaalibaba Endpoint terms
Cached input /M
Unknown
Context limit
131,072 tokens
Output limit
32,768 tokens
Tools
Listed by endpoint
Reasoning
Not listed by endpoint
Promotional discount
None reported
Uptime · last 30 minutes
99.02%
Input$0.150
Output$0.598
Precisionundeclared
Uptime (1d)92.35%
Seller
Venicevenice/fp8 Endpoint terms
Cached input /M
Unknown
Context limit
128,000 tokens
Output limit
16,384 tokens
Tools
Listed by endpoint
Reasoning
Not listed by endpoint
Promotional discount
None reported
Uptime · last 30 minutes
95.85%
Input$0.150
Output$0.750
Precisionfp8
Uptime (1d)95.29%
Seller
Nebiusnebius/fp8 Endpoint terms
Cached input /M
Unknown
Context limit
262,144 tokens
Output limit
235,929 tokens
Tools
Listed by endpoint
Reasoning
Not listed by endpoint
Promotional discount
None reported
Uptime · last 30 minutes
88.80%
Input$0.200
Output$0.600
Precisionfp8
Uptime (1d)87.82%
Seller
StreamLakestreamlake Endpoint terms
Cached input /M
Unknown
Context limit
128,000 tokens
Output limit
32,000 tokens
Tools
Listed by endpoint
Reasoning
Not listed by endpoint
Promotional discount
40% · already included in these rates
Uptime · last 30 minutes
99.61%
Input$0.210
Output$0.840
Precisionundeclared
Uptime (1d)99.40%
Seller
Googlegoogle-vertex/us-south1 Endpoint terms
Cached input /M
Unknown
Context limit
262,144 tokens
Output limit
16,384 tokens
Tools
Listed by endpoint
Reasoning
Not listed by endpoint
Promotional discount
None reported
Uptime · last 30 minutes
100.00%
Input$0.220
Output$0.880
Precisionundeclared
Uptime (1d)99.69%
Seller
Googlegoogle-vertex/us-south1 Endpoint terms
Cached input /M
Unknown
Context limit
262,144 tokens
Output limit
16,384 tokens
Tools
Not listed by endpoint
Reasoning
Not listed by endpoint
Promotional discount
None reported
Uptime · last 30 minutes
100.00%
Input$0.250
Output$1.00
Precisionundeclared
Uptime (1d)99.44%
Independent scores
LMArena Elo
1419.3default
LMArena Elo under CC BY 4.0. Absence of another board is not a score of zero.
Best rankings by task
Data visualisation#981078
3D#1031002
Websites#1051064
Code categories#1051044
UI components#110972
Game development#117964
Per-task Elo and rank from Design Arena, via OpenRouter’s model feed. The rank is the feed’s, and it counts models OpenRouter does not list.
Qwen3-235B-A22B-Instruct-2507 is a multilingual, instruction-tuned mixture-of-experts language model based on the Qwen3-235B architecture, with 22B active parameters per forward pass. It is optimized for general-purpose text generation, including instruction following,...
Questions this page answers
What does Qwen3 235B A22B Instruct 2507 cost?
$0.090/M in, $0.550/M out in this catalogue, as of the fetch date on this page. That is the API row, not a subscription.
Does Qwen3 235B A22B Instruct 2507 have an independent quality score?
Qwen3 235B A22B Instruct 2507 has an LMArena score in this catalogue.
What beats Qwen3 235B A22B Instruct 2507?
DeepSeek V4.1 Flash has an equal-or-higher measured score and an equal-or-lower price.
Does this page use Artificial Analysis scores for Qwen3 235B A22B Instruct 2507?
No. Artificial Analysis figures are not published here. Quality on this page is LMArena Elo where a score exists; otherwise the row is unrated.
Is the cheapest Qwen3 235B A22B Instruct 2507 endpoint the same product?
Headline cheapest is a lower precision. Like-for-like at the best declared precision is $0.090/M from DeepInfra.
MODEL MONUMENT
Save or share this proof
Preview and more sharing options
Qwen: Qwen3 235B A22B Instruct 2507. Balanced workload. Better value option available. Evidence as of 06 Oct 2026.
The verdict as one SVG, for a README or a docs page. It states this model's dominance
status at the balanced workload on the LMArena lens, with the date it was computed, and
it is rebuilt with the catalogue — so it changes when the verdict changes, including to
one you would rather it did not.