The same weights, different shops. Cheapest is not like-for-like when serving precision differs.
Headline cheapest is a lower precision. Like-for-like at the best declared precision is $0.090/M from SiliconFlow.
5 provider offers · rates, limits and conditions
Rates are USD per million tokens. Endpoint terms can differ even when the quoted price and precision match. A listed parameter is a provider declaration, not a task-success test.
Seller
Input
Output
Precision
Uptime (1d)
Seller
StreamLakestreamlake Endpoint terms
Cached input /M
Unknown
Context limit
128,000 tokens
Output limit
32,000 tokens
Tools
Listed by endpoint
Reasoning
Not listed by endpoint
Promotional discount
55% · already included in these rates
Uptime · last 30 minutes
99.60%
Input$0.048
Output$0.193
Precisionundeclared
Uptime (1d)99.12%
Seller
SiliconFlowsiliconflow/fp8 Endpoint terms
Cached input /M
Unknown
Context limit
262,144 tokens
Output limit
235,929 tokens
Tools
Listed by endpoint
Reasoning
Not listed by endpoint
Promotional discount
None reported
Uptime · last 30 minutes
90.47%
Input$0.090
Output$0.300
Precisionfp8
Uptime (1d)95.90%
Seller
DekaLLMdekallm Endpoint terms
Cached input /M
Unknown
Context limit
262,144 tokens
Output limit
235,929 tokens
Tools
Listed by endpoint
Reasoning
Not listed by endpoint
Promotional discount
None reported
Uptime · last 30 minutes
98.87%
Input$0.090
Output$0.300
Precisionundeclared
Uptime (1d)98.69%
Seller
Nebiusnebius/fp8 Endpoint terms
Cached input /M
Unknown
Context limit
262,144 tokens
Output limit
235,929 tokens
Tools
Listed by endpoint
Reasoning
Not listed by endpoint
Promotional discount
None reported
Uptime · last 30 minutes
93.30%
Input$0.100
Output$0.300
Precisionfp8
Uptime (1d)92.77%
Seller
Alibabaalibaba Endpoint terms
Cached input /M
Unknown
Context limit
131,072 tokens
Output limit
32,768 tokens
Tools
Listed by endpoint
Reasoning
Not listed by endpoint
Promotional discount
None reported
Uptime · last 30 minutes
99.92%
Input$0.130
Output$0.520
Precisionundeclared
Uptime (1d)99.88%
Independent scores
LMArena Elo
1383.6default
LMArena Elo under CC BY 4.0. Absence of another board is not a score of zero.
Qwen3-30B-A3B-Instruct-2507 is a 30.5B-parameter mixture-of-experts language model from Qwen, with 3.3B active parameters per inference. It operates in non-thinking mode and is designed for high-quality instruction following, multilingual understanding, and...
Questions this page answers
What does Qwen3 30B A3B Instruct 2507 cost?
$0.048/M in, $0.193/M out in this catalogue, as of the fetch date on this page. That is the API row, not a subscription.
Does Qwen3 30B A3B Instruct 2507 have an independent quality score?
Qwen3 30B A3B Instruct 2507 has an LMArena score in this catalogue.
What beats Qwen3 30B A3B Instruct 2507?
Nothing is both better and cheaper than Qwen3 30B A3B Instruct 2507.
Does this page use Artificial Analysis scores for Qwen3 30B A3B Instruct 2507?
No. Artificial Analysis figures are not published here. Quality on this page is LMArena Elo where a score exists; otherwise the row is unrated.
Is the cheapest Qwen3 30B A3B Instruct 2507 endpoint the same product?
Headline cheapest is a lower precision. Like-for-like at the best declared precision is $0.090/M from SiliconFlow.
MODEL MONUMENT
Save or share this proof
Preview and more sharing options
Qwen: Qwen3 30B A3B Instruct 2507. Balanced workload. Frontier. Evidence as of 23 Sept 2026.
The verdict as one SVG, for a README or a docs page. It states this model's dominance
status at the balanced workload on the LMArena lens, with the date it was computed, and
it is rebuilt with the catalogue — so it changes when the verdict changes, including to
one you would rather it did not.