Price from Alibaba: the cheapest offer that is serving, at standard delivery and at the maker’s precision or better, at its standard rate.
Every price dimension · USD per million tokens
Input
$0.390/M
Output
$2.34/M
Cached input
Unknown/M
Cache write
Unknown/M
Cache write 1h
Unknown/M
Reasoning
Unknown/M
Who sells it
The same weights, different shops. Cheapest is not like-for-like when serving precision differs.
Headline cheapest is a lower precision. Like-for-like at the best declared precision is $0.450/M from DeepInfra.
10 provider offers · rates, limits and conditions
Rates are USD per million tokens. Endpoint terms can differ even when the quoted price and precision match. A listed parameter is a provider declaration, not a task-success test.
Seller
Input
Output
Precision
Uptime (1d)
Seller
Alibabaalibaba Endpoint terms
Cached input /M
Unknown
Context limit
262,144 tokens
Output limit
65,536 tokens
Tools
Listed by endpoint
Reasoning
Listed by endpoint
Promotional discount
None reported
Uptime · last 30 minutes
100.00%
Input$0.390
Output$2.34
Precisionundeclared
Uptime (1d)99.87%
Seller
DeepInfradeepinfra/fp8 Endpoint terms
Cached input /M
$0.220
Context limit
262,144 tokens
Output limit
81,920 tokens
Tools
Listed by endpoint
Reasoning
Listed by endpoint
Promotional discount
None reported
Uptime · last 30 minutes
100.00%
Input$0.450
Output$3.00
Precisionfp8
Uptime (1d)99.62%
Seller
Parasailparasail/fp8 Endpoint terms
Cached input /M
$0.300
Context limit
262,144 tokens
Output limit
235,929 tokens
Tools
Listed by endpoint
Reasoning
Listed by endpoint
Promotional discount
None reported
Uptime · last 30 minutes
99.88%
Input$0.500
Output$3.60
Precisionfp8
Uptime (1d)99.84%
Seller
AtlasCloudatlas-cloud/fp8 Endpoint terms
Cached input /M
$0.550
Context limit
262,144 tokens
Output limit
65,536 tokens
Tools
Listed by endpoint
Reasoning
Listed by endpoint
Promotional discount
None reported
Uptime · last 30 minutes
99.07%
Input$0.550
Output$3.50
Precisionfp8
Uptime (1d)87.24%
Seller
DigitalOceandigitalocean Endpoint terms
Cached input /M
$0.110
Context limit
131,072 tokens
Output limit
117,964 tokens
Tools
Listed by endpoint
Reasoning
Listed by endpoint
Promotional discount
None reported
Uptime · last 30 minutes
98.13%
Input$0.550
Output$3.50
Precisionundeclared
Uptime (1d)96.62%
Seller
Phalaphala Endpoint terms
Cached input /M
$0.225
Context limit
262,144 tokens
Output limit
235,929 tokens
Tools
Listed by endpoint
Reasoning
Listed by endpoint
Promotional discount
None reported
Uptime · last 30 minutes
99.76%
Input$0.550
Output$3.50
Precisionundeclared
Uptime (1d)99.59%
Seller
GMICloudgmicloud/fp8 Endpoint terms
Cached input /M
Unknown
Context limit
262,144 tokens
Output limit
235,929 tokens
Tools
Listed by endpoint
Reasoning
Listed by endpoint
Promotional discount
None reported
Uptime · last 30 minutes
Unknown
Input$0.600
Output$3.60
Precisionfp8
Uptime (1d)78.05%
Seller
StreamLakestreamlake Endpoint terms
Cached input /M
$0.120
Context limit
256,000 tokens
Output limit
64,000 tokens
Tools
Not listed by endpoint
Reasoning
Listed by endpoint
Promotional discount
None reported
Uptime · last 30 minutes
Unknown
Input$0.600
Output$3.60
Precisionundeclared
Uptime (1d)95.90%
Seller
Novitanovita Endpoint terms
Cached input /M
Unknown
Context limit
262,144 tokens
Output limit
65,536 tokens
Tools
Listed by endpoint
Reasoning
Listed by endpoint
Promotional discount
None reported
Uptime · last 30 minutes
99.18%
Input$0.600
Output$3.60
Precisionundeclared
Uptime (1d)99.08%
Seller
Venicevenice Endpoint terms
Cached input /M
Unknown
Context limit
128,000 tokens
Output limit
32,768 tokens
Tools
Listed by endpoint
Reasoning
Listed by endpoint
Promotional discount
None reported
Uptime · last 30 minutes
Unknown
Input$0.750
Output$4.50
Precisionundeclared
Uptime (1d)86.05%
Independent scores
LMArena Elo
1438default
LMArena Elo under CC BY 4.0. Absence of another board is not a score of zero.
Best rankings by task
SVG#491139
3D#551169
Code categories#581184
Websites#591197
Data visualisation#601182
Game development#671153
UI components#681170
Per-task Elo and rank from Design Arena, via OpenRouter’s model feed. The rank is the feed’s, and it counts models OpenRouter does not list.
The Qwen3.5 series 397B-A17B native vision-language model is built on a hybrid architecture that integrates a linear attention mechanism with a sparse mixture-of-experts model, achieving higher inference efficiency. It delivers...
Questions this page answers
What does Qwen3.5 397B A17B cost?
$0.390/M in, $2.34/M out in this catalogue, as of the fetch date on this page. That is the API row, not a subscription.
Does Qwen3.5 397B A17B have an independent quality score?
Qwen3.5 397B A17B has an LMArena score in this catalogue.
What beats Qwen3.5 397B A17B?
MiMo-V2.6-Pro has an equal-or-higher measured score and an equal-or-lower price.
Does this page use Artificial Analysis scores for Qwen3.5 397B A17B?
No. Artificial Analysis figures are not published here. Quality on this page is LMArena Elo where a score exists; otherwise the row is unrated.
Is the cheapest Qwen3.5 397B A17B endpoint the same product?
Headline cheapest is a lower precision. Like-for-like at the best declared precision is $0.450/M from DeepInfra.
MODEL MONUMENT
Save or share this proof
Preview and more sharing options
Qwen: Qwen3.5 397B A17B. Balanced workload. Better value option available. Evidence as of 06 Oct 2026.
The verdict as one SVG, for a README or a docs page. It states this model's dominance
status at the balanced workload on the LMArena lens, with the date it was computed, and
it is rebuilt with the catalogue — so it changes when the verdict changes, including to
one you would rather it did not.